NEET UG 2026 was held on 21 June. By evening, the NTA had published the official question paper PDF on its website. We downloaded it, handed it to an AI agent, and asked it to sit the exam.
No answer key. No prior exposure to the paper. No special NEET preparation. Just the scanned PDF, the ShunyaSaarthi NCERT knowledge base, and a reasoning model told to answer every question and show its work.
What follows is the honest result — not a curated demo, but the actual benchmark with the errors left in.
Contents
Section 01 · The Setup
The Setup
NEET UG 2026 Set 50 "SUSHRUT" is a 180-question paper across Physics, Chemistry, and Biology. Scoring is +4 for a correct answer, −1 for a wrong one, and 0 for skipped. Maximum marks: 720.
The NTA publishes question booklets as image-scanned PDFs — not text-extractable files, not digital test interfaces. The agent had to read the same document a student receives in the exam hall.
The knowledge base it drew on: ShunyaSaarthi — a CBSE AI tutoring platform built on NCERT Class XII content for Physics, Chemistry, and Biology. 2,563 Physics XII semantic chunks, 969 Biology XII points, Chemistry XII fully indexed. The system was built for tutoring, not for exam-taking. We gave it a live exam.
One constraint worth flagging upfront: NTA embeds answer keys inside some question booklets but not all. Of the 180 questions in Set 50, 105 had keys embedded inline. The remaining 75 — covering large blocks of Physics and Chemistry — require a separate official answer PDF that NTA had not yet released at benchmark time. We scored against what we could verify. The agent answered all 180; 75 results are banked for reconciliation.
Section 02 · The Pipeline
The Pipeline
Four stages, no human in the loop between them.
Vision Extraction
GPT-4o reads each scanned page as a high-resolution image at 150 DPI. It extracts question text, option text, flags diagram questions, and captures any embedded answer key — handling circuit diagrams, ray optics figures, truth tables, and match-list tables that defeat OCR entirely.
NCERT Retrieval (RAG)
The question stem goes into the ShunyaSaarthi knowledge base. Top-3 semantically matching NCERT chunks are retrieved using BM25 + dense vector search with RRF hybrid re-ranking across the subject's Qdrant collection.
Grounded Reasoning
o4-mini at temperature=0 receives the question, the four options, and the retrieved NCERT context. It reasons step-by-step — deriving formulas, applying laws, eliminating options — then commits to exactly one answer. It never skips.
NEET Scoring
Agent answers are scored against the embedded NTA key at +4/−1. Questions without an embedded key are banked. Score is computed only on verified questions — no imputed marks, no assumptions.
Section 03 · The Results
The Results
105 keyed questions. 75 correct. 30 wrong. 0 skipped. Score: 270/420. Accuracy: 71.4%.
By subject, the picture is more nuanced:
| Subject | Keyed Q | Correct | Accuracy | NEET Marks | Pending |
|---|---|---|---|---|---|
| Biology | 75 | 62 | 83% | +248 − 13 = +235 | 15 Q |
| Chemistry | 16 | 10 | 62% | +40 − 6 = +34 | 29 Q |
| Physics | 14 | 3 | 21% | +12 − 11 = +1 | 31 Q |
| Total | 105 | 75 | 71.4% | +270 | 75 Q |
Biology is the headline. 83% accuracy on 75 keyed questions — from a knowledge base built for tutoring, not for exam preparation. The NCERT XII Biology coverage is deep enough that the agent outperforms the national median on that subject.
Physics is the asterisk. Only 14 of 45 Physics questions had embedded keys in Set 50 — and those 14 skewed toward numerically complex questions (Q1–9: variable friction, mean free path, inductor circuits, harmonic standing waves) and a set of optics/electrostatics problems (Q41–45) that were among the harder questions in the paper. 31 Physics questions are banked. The agent's 2024 Physics accuracy was 75% — we'll know the real 2026 Physics number when NTA releases the key.
The projected full-paper score, using subject-wise accuracy extrapolated to all 180 questions: ~505 / 720.
For context: NEET 2026 qualifying cutoff for General category was approximately 138/720. The projected 505 places the agent well into the competitive range — not top-100, but solidly competitive for state-level medical admissions.
We also ran the same agent on NEET 2024 Set Q1 — a fully keyed Aakash mirror PDF with clean extractable text. That run: 61/80 correct, 76.2% accuracy, 225/320 marks. Biology improved year-over-year (83% vs 78%). The overall accuracy dip from 2024 to 2026 is partly paper source (vision extraction vs text extraction) and partly the composition of which questions happened to have embedded keys.
Section 04 · How the Agent Actually Thinks
How the Agent Actually Thinks
The most useful part of this benchmark is not the score — it's the reasoning trail. Every answer has a written justification. Here are three representative ones.
Q7 — Nuclear Reactions (Correct)
The question gave atomic masses and asked for the Q-value of a uranium alpha decay. The agent retrieved the NCERT formula, computed the mass defect (0.004 u), multiplied by 931.5 MeV/u, and got 3.726 MeV. It matched the key. The retrieval brought the exact NCERT passage; the reasoning was mechanical from there.
Q47 — Restriction Endonucleases (Correct)
A four-option question on the properties of restriction enzymes. The agent retrieved Biology XII Chapter 11, identified the passage on palindromic sequence recognition and sticky-end generation, eliminated the distractor options by name (ligase is not an endonuclease), and committed. Clean, traceable, no hallucination.
Q3 — Inductor Circuit Diagram (Wrong)
Two inductor configurations labelled P and Q in a scanned diagram. The agent extracted the image correctly, inferred the topology — but swapped the P/Q labels. It reasoned correctly from the wrong premise: parallel gives L/2, series gives 2L, ratio is 1/4. The key says 1/2. The physics reasoning was sound. The vision extraction was the failure point.
This pattern — correct reasoning on wrong extracted input — accounts for roughly 12% of errors. It's a vision problem, not a knowledge problem.
Section 05 · Where It Failed — and Why
Where It Failed — and Why
30 wrong answers. We categorised every one.
KB Gaps — 13 errors (44%)
Topics where the NCERT XII knowledge base returned irrelevant chunks and the agent reasoned cold: polyatomic gas degrees of freedom, Bt toxin alkaline activation, oxalate complex chirality, gymnosperm gametophyte independence, organic compound polarity ordering. All addressable by enriching the KB.
Option-Map Artefacts — 6 errors (20%)
In Assertion-Reason and match-list questions, GPT-4o sometimes renumbers options differently from the printed booklet. The agent's reasoning was correct in at least 4 of these 6 cases — only the letter commitment was wrong. An extraction fix, not a knowledge fix.
Calculation Errors — 6 errors (20%)
Mean free path ratios, titration volumes, banked circular track angles, stopping potential. The formulas were retrieved correctly; the arithmetic was wrong. Consistent with o4-mini's known behaviour on multi-step numerical problems at temperature=0.
Diagram Misreads — 3 errors (10%)
Circuit topology swaps, p-n junction orientation ambiguity, orbital overlap spatial depth. GPT-4o correctly extracted ~85% of diagram questions. Three fell through where the image resolution or spatial complexity defeated confident interpretation.
The split matters for what to fix next. KB gaps → expand the knowledge base into NCERT XI and edge-case XII topics. Option-map artefacts → better extraction post-processing. Calculation errors → a verification pass or code interpreter tool. Diagram misreads → higher DPI rendering or a specialised visual reasoning step.
Section 06 · What This Tells Us
What This Tells Us
Three things that surprised us, and one that didn't.
Biology NCERT coverage is real. 83% blind accuracy on a live national exam is not a fluke. The ShunyaSaarthi knowledge base covers NCERT XII Biology deeply enough that the agent can answer questions it has never seen, on topics it was not specifically prepared for. That's the retrieval working as designed.
Vision extraction is the current ceiling. The agent's knowledge and reasoning are ahead of its ability to read the input. A higher-quality extraction layer — better DPI, better post-processing for option numbering, a dedicated visual QA pass for circuit/diagram questions — would likely add 5–8 percentage points without touching the knowledge base at all.
The agent never skips. Of 180 questions, it committed to an answer on every single one. That is both its strength (no negative marks from blank answers) and its risk (wrong answers cost marks; confident wrong answers cost the same as uncertain ones). A calibration layer — "I am not confident enough to commit here" — would improve the penalty-adjusted score.
71.4% blind accuracy on a national medical entrance exam is a starting point, not an endpoint. The knowledge base was built for tutoring CBSE students. It was not fine-tuned on PYQs, not optimised for NEET question styles, not run with a NEET-specific prompt. The 505/720 projection comes from a system designed to explain electrostatics to a Class XII student — not from a system designed to maximise NEET marks.
That distinction matters because there's a clear path to 85%+ from here — and it doesn't require a bigger model. It requires better knowledge, better extraction, and better calibration.
The 75 pending questions will be reconciled against the NTA official key when it is released. Full benchmark report with question-level reasoning and error analysis: agentadda.in/reports/neet-2026
Benchmark simulated Sep 9, 2026. Exam date: 21 June 2026. NTA official Set 50 "SUSHRUT" PDF. Scoring: +4 correct / −1 wrong / 0 skipped. Models: GPT-4o (vision extraction) + o4-mini temperature=0 (reasoning). KB: ShunyaSaarthi NCERT XII — Physics, Chemistry, Biology.