How to Prepare for Quant Interviews Using AI (ChatGPT, Claude, Gemini)
What LLMs are genuinely good at in quant interview prep, where the research says they fail, the one place you must not use them, and how to connect ChatGPT or Claude straight to a graded problem bank over MCP.

Every quant candidate I talk to now has ChatGPT, Claude, or Gemini open in a tab while they study. Almost none of them are using it well. The typical pattern is to paste a probability question in, read the answer, feel like you understood it, and move on. Two weeks later you cannot reproduce the argument in front of an interviewer, which is the only place it counts.
That is a shame, because a modern LLM is genuinely the best study partner that has ever existed for this particular exam. It is patient, it is available at 1am, it will explain the same conditional expectation five different ways, and it never sighs. The trick is knowing which parts of your prep it should touch and which parts it will quietly ruin.
This is what actually works, what the research says about where these models break, the one place you must not use them, and how to wire a model directly into a structured problem bank so your prep stops being a pile of open tabs.
What LLMs are genuinely good at
Concept building from wherever you are. The hardest thing about quant prep is that the standard resources assume a starting point you may not have. Blitzstein assumes measure-free comfort with random variables. The green book assumes you already think in expectations. A model does not assume anything. You can say "explain the law of total expectation to someone who understands averages but has never seen conditioning on a random variable" and get a version pitched at you, then ask for it again with a concrete dice example, then again as a one-line intuition you can hold in your head during an interview. That ladder, from formal to concrete to compressed, is the whole game in concept building, and it used to require a patient human.
Translating between resources. Nobody prepares from one book. You have Zhou for brainteasers, a stats course for estimators, a firm's own guide for market making, and a forum thread for the actual question someone got asked. These use different notation and different framings for the same object. Paste two of them in and ask "are these the same concept, and where do the notations disagree" and you save an evening of confusion. This is a pure translation task, which is exactly what language models are built for.
Generating variants. Interviewers change the numbers. Memorizing that the answer to "expected flips until two heads in a row" is 6 gets you nothing when they ask for three heads, or for heads-tails, or for a biased coin. Ask the model to take a problem you just solved and produce four variants that break your solution method. Then solve those. This is the single highest-leverage prompt in quant prep and almost nobody runs it.
Mock interviews with a real interviewer's manner. Give the model a role and constraints: senior trader at a market making firm, ask one problem at a time, never give the answer until I commit to one, hint only when I ask, and press me on why after every answer. Then talk out loud and type what you would have said. The value is not the problem. The value is practising the narration, because quant interviews are graded on how you reason aloud far more than on whether you land the number.
Post-mortems. After a real interview, write down every question you can remember and hand it to the model with what you actually said. Ask it to find the specific step where your reasoning went wrong, not to re-solve the problem. Candidates consistently misdiagnose their own failures as "I was nervous" when the real problem was a setup error they repeat every time.
Where they quietly fail
Here is the part that matters, and it is not the part people worry about.
Models are good at spotting errors and bad at confirming correctness. Research on mathematical hallucination detection keeps finding the same asymmetry: when a solution contains a mistake, detection rates run above 90%, but when a solution is correct, models validate it reliably only a fraction of the time. The practical consequence is brutal. Ask "is my answer right?" and you will get a confident yes far more often than you have earned one, and you will also get spurious objections to correct work. The model is not a grader. It is a critic with a bad calibration.
They also cannot tell you a question is broken. A benchmark released this year that specifically tests whether models will refuse an ill-posed problem found that no frontier model cleared 50% on it. A related result on unsolvable math problems found that models often notice something is wrong mid-reasoning and then fabricate a derivation anyway, because they are trained to produce an answer. If you transcribe a brainteaser slightly wrong, which happens constantly when you are working from memory or a forum post, the model will not tell you. It will confidently solve the wrong problem.
And a correct final answer does not mean correct reasoning. One process-level evaluation this year found that roughly 7% of responses with the right final answer had reasoning the graders scored as failures. In an interview, that 7% is 100% of your outcome, because the interviewer is grading the reasoning.
So the rule is: use the model to explain, generate, and challenge. Use something else to grade. The moment your prep depends on the model being right about whether you were right, it stops being prep and becomes a confidence-inflation machine.
The prompt patterns that survive contact
Four that are worth saving as reusable prompts.
- The tutor. "Explain X. Then give me the two-sentence version. Then ask me one question that tests whether I actually got it." The closing question is the part everyone leaves out and the part that does the work.
- The adversary. "Here is my solution. Do not tell me whether it is right. List every assumption I made and name the one most likely to be wrong." Asking for assumptions instead of a verdict routes around the calibration problem entirely.
- The variant generator. "Change one structural feature of this problem so my method breaks. Do not tell me which feature."
- The interviewer. Role, one question at a time, no answers until I commit, follow-up on every response. Run it timed.
Notice that none of them ask the model to be the source of truth.
Applications, resumes, and firm research
Away from the math, the picture is simpler and the tools are unambiguously useful.
Resume rewriting for a specific job description works well, especially the mechanical part: matching the vocabulary the posting uses, cutting a bullet to one line, quantifying an outcome you described vaguely. Firm research works well too. Ask for the structure of a given firm's loop, what each round is testing, and what a strong answer sounds like, then verify the specifics against the firm's own careers page, because this is exactly the kind of detail models get subtly out of date on.
Cover letters and networking messages are fine to draft with a model and terrible to send unedited. Recruiters at these firms read hundreds a week and the generated register is obvious. Use the draft for structure, then rewrite every sentence in your own voice.
One honest note on the engineering side. Experienced developers at trading firms use these tools heavily and say so. The trap for candidates is that if you let a model write your projects, you cannot defend the architecture when someone asks why you chose a lock-free queue over a mutex. Do the design yourself. Use the model for refactors, for explaining an error, and for reviewing what you wrote.
The one place you must not use AI
During the assessment itself.
This changed fast. Assessment vendors now ship AI-assistance detection alongside screen lockdown, tab-switch logging, and webcam analysis, and they publish their detection results. CodeSignal reported that flagged cheating attempts on proctored assessments more than doubled over 2025. Hudson River Trading lists "cheating with LLMs" explicitly among its rejection signals and says it detects AI-generated code patterns. Two Sigma prohibits AI use on its online assessment outright. Firms that do publish candidate guidelines, like SEI's AI ethics policy for candidates, draw the line in the same place: AI is fine for preparing, researching, and drafting your resume, and disqualifying when it develops your answers.
Assume detection works, assume the policy is no unless the invitation explicitly says yes, and take the assessment on a clean machine. The asymmetry is obvious. You are risking a candidacy at a firm that pays more than almost any other employer on earth to save twenty minutes on a screening round.
The missing piece: your chat does not know you
Even if you do all of the above well, there is a structural gap. The model has no idea what you have already solved, which topics you keep avoiding, whether your answer was actually correct, or which firm you are interviewing with next week. Every session starts from nothing. You end up narrating your own progress back to it, badly, and the recommendations it gives you are generic because the input was generic.
That gap is what the Model Context Protocol closes. MCP is an open standard, supported across Claude, ChatGPT, Gemini, Cursor, VS Code and others, that lets a chat client connect to an external service and call its tools directly. The chat stops being a closed room and becomes a front end to a system that holds real state.
PuzzledQuant over MCP
PuzzledQuant is, as far as we know, the only quant interview prep platform that ships an MCP server. Connect it once and your assistant can work against the real problem bank and your real progress, inside whichever chat app you already live in.
What that gives you concretely:
- Practice with grading that is not the model's opinion.
get_question_of_the_dayorstart_problem_playgroundopens a tracked session; you reason it through with the model as your tutor, thensubmit_answergrades it against the stored answer. The verdict comes from the platform, not from the assistant's guess. That single split fixes the calibration problem described above: the model explains, the platform judges. - Recommendations based on what you have actually done.
get_recommended_problemspicks from the topics where you have solved the smallest share of what is available, so the thing you have been avoiding is the thing that surfaces. - Company-targeted prep.
search_problemsfilters by company, topic, difficulty, unsolved, and bookmarked. "Give me the Optiver probability problems I have not solved yet" is a sentence that now does something. - Real progress, not vibes.
get_progress_summaryandget_leaderboard_rankgive the assistant your solved count, your attempting count, and where you sit, so a weekly plan is built on data. - Sets and bookmarks from the chat.
bookmark_problemandupdate_problem_setlet you file a problem you fumbled into a "revisit before Jane Street" set without leaving the conversation. - Jobs in the same loop.
search_jobsandget_job_detailssearch current quant openings with filters for company, country, position type, and salary, so the "what am I preparing for" question and the "who is hiring" question live in one place.
The whole thing runs over a remote HTTP endpoint with OAuth, so there is nothing to install and nothing to keep running on your laptop. For Claude Code it is one line:
claude mcp add --transport http puzzledquant https://www.puzzledquant.com/mcpFor ChatGPT, Claude Desktop and Web, Cursor, VS Code, Gemini, Codex, and OpenClaw, the copy-paste config for each is on the MCP setup guide. It takes about a minute.
Putting it together
A week that uses all of this without eating your life:
- Daily, 10 minutes. Mental arithmetic drills. No AI involved; this is pure reflex training.
- Daily, 40 minutes. Ask your assistant for a recommended problem over MCP, solve it yourself first, then use the model as a tutor on the parts you fumbled, then submit for a real grade. Bookmark anything you got for the wrong reason.
- Twice a week, 20 minutes. Take one problem you solved and have the model generate variants that break your method. Solve those.
- Once a week, 45 minutes. Timed mock interview with the model in the interviewer role, out loud, on company-tagged problems for wherever you are actually applying.
- Once a week, 15 minutes. Ask for a progress summary and let the model tell you which topic you have been dodging. You already know. Make it say it anyway.
The tools have genuinely changed what a single motivated candidate can do alone. What has not changed is that the interview measures whether you can reason out loud under pressure, and the only way to build that is repetition against problems that grade you honestly.
Start with the problem set, then connect the MCP server so your assistant knows what you have done.