2.1Why a general AI model fails the serious candidate.
A model trained to know a little about everything is the wrong instrument for an examination that demands precision in one thing. Any general purpose model is rewarded, during training, for sounding plausible across a wide range of topics. That is a sensible objective for a curious reader. It is a poor one for a candidate who must, in a few months, sit a paper that will partly decide the next four years of their life.
Two failures follow. Breadth dilutes accuracy: a general model produces respectable answers on most subjects and rarely outstanding ones on any, while an examination rewards the outstanding answer and is indifferent to the merely respectable. And when a general model is uncertain, it invents. The most comprehensive recent survey of this failure, by Ji and others, finds it persistent across model sizes and architectures.4 Bender and others argue that scaling alone does not eliminate it, since the training objective continues to reward fluency over warrantability.5 Bommasani and others add that breadth is bought at the cost of precision in any one domain.6
The narrow alternative is well established. Howard and Ruder showed that a general model can be adapted to a target subject by further training on a small, well-chosen library, with substantial gains in accuracy within it.7 Lewis and others showed that letting a model look an answer up in a curated library, before producing it, sharply reduces fabrication.8
We have applied both. Vault 720 AI is trained on the annotated study materials and worked problems of candidates who scored 720 on NEET; Vault 45 AI on examiner-marked top-band papers, mark schemes and tutorial corrections. Both are curated under faculty supervision and instructed to decline, where uncertain, rather than invent. On the public benchmarks by which a general model is measured, both lose. On the mark a candidate earns on a real paper marked by a real examiner, they win. Where the purpose is to know a few things deeply, breadth is a vice.
- 4.Z Ji and others, 'Survey of Hallucination in Natural Language Generation' (2023) 55(12) ACM Computing Surveys 1. The survey is the most comprehensive recent treatment of confident error in language models, and is the source upon which we principally rely for the empirical claim.
- 5.EM Bender, T Gebru, A McMillan-Major and S Shmitchell, 'On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?' (Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021) 610.
- 6.R Bommasani and others, 'On the Opportunities and Risks of Foundation Models' (Stanford Center for Research on Foundation Models 2021). We are aware that the framing of "foundation models" has been criticised; we cite the work for the breadth-versus-precision claim, which is, in our view, well supported in §§4-5 of the report regardless of the framing.
- 7.J Howard and S Ruder, 'Universal Language Model Fine-Tuning for Text Classification' (Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018) 328.
- 8.P Lewis and others, 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks' (2020) 33 Advances in Neural Information Processing Systems 9459.