vaultedcorp
Back
vaultedcorp
The Vaulted Review
Volume 01 · Number 01 · Inaugural Edition

Plausibility is not knowledge.

Abstract. The dominant assumption of the present moment — that a sufficiently large general purpose AI model is the right tool for any educational task — is wrong in the case of preparing a serious candidate for a serious examination. Examinations reward precision in a small set of moves; a general model is trained for plausibility across many. We set out the case for a deliberately narrow alternative: four AI systems, each trained on the recorded practice of candidates who have already reached the ceiling of a specific examination — the IB Diploma, JEE Main, CBSE Class 12 and NEET — and designed to refuse where they have not been trained.

1Approach

Each of the four examinations we cover rewards a specific and learnable set of patterns: the answer that earns the maximum mark in IB Higher Level Economics, the method that solves a JEE Main Physics problem inside the required time, the structure that separates the CBSE top-scorer's response from the rest of the distribution, the precision that carries a NEET candidate to the top. These patterns are not theoretical. They are documented in the work of candidates who have already scored at the top. We train a language model narrowly on that work.

The practical effect is that a model trained on the annotated study materials of top-scoring NEET candidates will, placed before a NEET problem, respond the way a top scorer would. A model trained only on the general web will not. The open question is not whether a candidate who trains on the right material outperforms one who does not — they will — but whether that material can be made available at scale, in a form a serious candidate can act upon.

We should be plain about what we are not offering. Not content: the internet is drowning in it, and the candidates who matter to us have read more than they need. Not motivation: they have already decided they want to be exceptional, and what they need is leverage, not encouragement. Not a shortcut: we have not found a real one, and we have stopped looking.

The Vaulted Review1

2Principle

We do not teach. We engineer outcomes. What follows is the principle by which we decide what to undertake, what to refuse, and with whom we are willing to work.

2.1Why a general AI model fails the serious candidate.

A model trained to know a little about everything is the wrong instrument for an examination that demands precision in one thing. Any general purpose model is rewarded, during training, for sounding plausible across a wide range of topics. That is a sensible objective for a curious reader. It is a poor one for a candidate who must, in a few months, sit a paper that will partly decide the next four years of their life.

Two failures follow. Breadth dilutes accuracy: a general model produces respectable answers on most subjects and rarely outstanding ones on any, while an examination rewards the outstanding answer and is indifferent to the merely respectable. And when a general model is uncertain, it invents. The most comprehensive recent survey of this failure, by Ji and others, finds it persistent across model sizes and architectures.4 Bender and others argue that scaling alone does not eliminate it, since the training objective continues to reward fluency over warrantability.5 Bommasani and others add that breadth is bought at the cost of precision in any one domain.6

The narrow alternative is well established. Howard and Ruder showed that a general model can be adapted to a target subject by further training on a small, well-chosen library, with substantial gains in accuracy within it.7 Lewis and others showed that letting a model look an answer up in a curated library, before producing it, sharply reduces fabrication.8

We have applied both. Vault 720 AI is trained on the annotated study materials and worked problems of candidates who scored 720 on NEET; Vault 45 AI on examiner-marked top-band papers, mark schemes and tutorial corrections. Both are curated under faculty supervision and instructed to decline, where uncertain, rather than invent. On the public benchmarks by which a general model is measured, both lose. On the mark a candidate earns on a real paper marked by a real examiner, they win. Where the purpose is to know a few things deeply, breadth is a vice.

  1. 4.Z Ji and others, 'Survey of Hallucination in Natural Language Generation' (2023) 55(12) ACM Computing Surveys 1. The survey is the most comprehensive recent treatment of confident error in language models, and is the source upon which we principally rely for the empirical claim.
  2. 5.EM Bender, T Gebru, A McMillan-Major and S Shmitchell, 'On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?' (Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021) 610.
  3. 6.R Bommasani and others, 'On the Opportunities and Risks of Foundation Models' (Stanford Center for Research on Foundation Models 2021). We are aware that the framing of "foundation models" has been criticised; we cite the work for the breadth-versus-precision claim, which is, in our view, well supported in §§4-5 of the report regardless of the framing.
  4. 7.J Howard and S Ruder, 'Universal Language Model Fine-Tuning for Text Classification' (Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018) 328.
  5. 8.P Lewis and others, 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks' (2020) 33 Advances in Neural Information Processing Systems 9459.
The Vaulted Review2

2.2Why the practice a model trains on shapes the model it becomes.

The quality of the practice a model trains on shapes the quality of what it produces. This applies with particular force to examination preparation, where the gap between top-scorer practice and average preparation is wide and consistently underestimated.

Vygotsky identified the zone of proximal development: the gap between what a learner can do alone and what becomes possible within a more capable environment.9 Lave and Wenger extended this into the proposition that serious learning happens through sustained engagement with the practices of a community already competent at the thing being learned.10 Applied to an AI system the principle is direct. A model trained on the compiled practice of top scorers is situated within a community of top-level performance; a model trained on the general web is not. The quality of the corpus is the quality of the ceiling.

We built Vaultedcorp on that order of operations. Every training corpus in the suite begins with the work of candidates who have already reached the ceiling of the examination in question: annotated revision materials, worked problems, the notes made in the margins of difficult papers. The candidate who works with one of our models is practising alongside the compiled method of people who have already cleared the same bar. We add products only when the corpus meets that standard, and not before.

  1. 9.LS Vygotsky, Mind in Society: The Development of Higher Psychological Processes (Harvard University Press 1978). The essays collected in this volume contain the canonical statement of the zone of proximal development, particularly in chapter 6.
  2. 10.J Lave and E Wenger, Situated Learning: Legitimate Peripheral Participation (Cambridge University Press 1991).
The Vaulted Review3

2.3Why restraint is the most expensive of the disciplines.

We will not ship a piece of work until it is verifiably better than what already exists for the same purpose.

Whitehead observed nearly a century ago that the chief enemy of a serious education is what he called inert ideas: content received but never exercised, capabilities listed but never used.11 The contemporary form is the over-built educational technology product whose inventory of features exceeds the capacity of any one candidate to use them. Such products do not teach so much as overwhelm. We have erred, deliberately, in the opposite direction.

We test the judgement against candidates first, against marked papers second, and third against practitioners who have nothing to gain by being polite about our work. The greater part of what we build is retired before any candidate sees it. This is a costly way to build a company: a faster one would ship the same work earlier and grow at a higher rate. We would rather build a smaller company that means something than a larger one that does not.

  1. 11.AN Whitehead, The Aims of Education and Other Essays (Macmillan 1929), particularly the title essay's discussion of inert ideas at 1-23.
The Vaulted Review4

2.4Who we are actually serving, and why we refuse the rest.

Vaultedcorp exists for a particular kind of candidate and is not the right instrument for any other. The practices that raise examination performance are demanding, and they are not mysterious. There are essentially two: deliberate practice against a clear external standard, and prompt, substantive feedback calibrated to that standard.

Ericsson, Krampe and Tesch-Römer established that the difference between expert and ordinary performance, in domains as different as music and chess, is principally explained by the quantity and quality of practice undertaken against external feedback.12 Marton and Säljö distinguished the surface approach, which treats material as content to be memorised, from the deep approach, which treats it as something to be understood; the deep approach is reliably produced by feedback and by stakes of the kind a real external standard provides.13

We therefore treat our candidates as adults. We do not gamify their preparation, interrupt them with notifications, or flatter them by suggesting their hard subjects are easier than they are. We assume they will read with attention, sit with a difficult idea longer than is comfortable, and act on feedback rather than argue with it.

It follows that the median candidate is not our reader. We exist for the small population who have already decided that average is unacceptable and are now in search of leverage: typically ambitious, frequently anxious, and almost without exception reading more widely than their teachers expect. We serve that population, and decline to pretend we are useful to any other.

  1. 12.KA Ericsson, RT Krampe and C Tesch-Römer, 'The Role of Deliberate Practice in the Acquisition of Expert Performance' (1993) 100(3) Psychological Review 363.
  2. 13.F Marton and R Säljö, 'On Qualitative Differences in Learning: I. Outcome and Process' (1976) 46(1) British Journal of Educational Psychology 4. The deep-versus-surface distinction has been refined many times in the four decades since; the original paper remains the clearest statement of it.
The Vaulted Review5

3The Suite

The suite consists of four products, one for each of the IB Diploma, JEE Main, CBSE Class 12 and NEET. We built four instruments rather than one because the examinations are genuinely distinct, and a single instrument cannot serve them equally well. Each takes the recorded practice of candidates who have already reached the top of its examination, makes that practice legible, and gives a serious candidate access to it. The technical case is at §2.1.

All four present through the same seven modes. War Room turns mock results and remaining exam dates into a study plan that updates as scores and hours change. Exam returns feedback on a draft against the rubric the real examiner applies. Socratic walks through a problem step by step, declining to give the answer before the reasoning has been traced. Rapid is a feed of short, high-yield patterns that top scorers use and standard preparation overlooks. Flashcard generates spaced-repetition cards from any worked topic. Focus locks the candidate into one subject for a fixed block. Battle lets candidates challenge one another and climb leaderboards. The training data changes from product to product; the approach does not.

3.1Vault 45 AI.

Vault 45 AI is trained on the work that earns top marks in the International Baccalaureate Diploma, and only on that. The corpus comes from sources we are at liberty to use: notes, essays and revision documents contributed by alumni who finished in the top band and gave us permission to learn from their work; tutorial annotations written in the margins of student writing over years of supervising it; and worked answers and commentaries authored for this project. We hold no examination board's papers.

Place a Paper 1 economics question before it and it reasons the way a strong IB candidate would, because that is what its training material looks like. Ask it about a subject it has not been trained on and it says so. That last property is the one that most clearly separates it from a general chatbot, which will produce a confident answer even when it does not have one. In a curious reader's hands that is a curiosity; in a candidate's hands, the night before a paper, it is a hazard. The May 2024 IB session enrolled roughly one hundred and ninety thousand candidates, a small fraction of whom achieved the maximum mark of 7 in any given Higher Level subject.15 Vault 45 AI is built for the cohort whose ambition is that mark in the few subjects their application turns on.

Entrance mode is available as an optional add-on at $49 for three months. It addresses competitive university admissions tests and is built on the strategies of applicants who scored in the highest percentiles of those papers.

Access the product · vault45.ai
  1. 15.International Baccalaureate Organization, Diploma Programme Statistical Bulletin: May 2024 Examination Session (IBO 2024). The proportion of candidates achieving the maximum mark of 7 varies considerably by subject; in most Higher Level subjects it sits in the high single digits.
The Vaulted Review6

3.2Vault 300 AI.

Vault 300 AI is trained on the notes, revision materials and worked problems of top-scoring candidates on the Joint Entrance Examination Main, and only on those. JEE Main is the gateway to India's premier technical institutions and rewards a specific precision: applying a large body of conceptual and quantitative knowledge quickly, accurately, and under competitive pressure. The corpus comes from verified top scorers who gave explicit permission to train on their work, together with commentaries from people with direct experience of the examination's marking patterns.

JEE Main is taken by over one million candidates each year, and a top score is reached by a handful.16 We built Vault 300 AI because we had access to the working methods of candidates who reached it, and because the gap between their practice and the preparation available to everyone else seemed to us both large and reducible. A candidate working through a difficult integration or a multi-step reaction mechanism is not served by a confident wrong answer, so the model declines rather than invents.

Access the product · vault300.ai
  1. 16.National Testing Agency, Joint Entrance Examination Main: Annual Statistical Report (NTA 2024). Registered candidates for the 2024 session exceeded one million across both sessions of the examination. The top score is achieved by a very small fraction of the candidate population in any given session.
The Vaulted Review7

3.3Vault 500 AI.

Vault 500 AI is trained on the study materials, past paper responses and revision notes of top-scoring candidates in the Central Board of Secondary Education Class 12 examinations, and only on those. A top score there is the qualifying standard for India's most selective undergraduate programmes and a prerequisite for many competitive entrance examinations. The corpus is drawn from candidates who cleared the papers in the top fraction of the distribution, with commentaries from people who know the board's marking standards.

CBSE Class 12 is sat by over fourteen million candidates a year, one of the largest single-cycle school-leaving examinations in the world.17 We built Vault 500 AI because the examination is consequential, the competition for top marks is enormous, and the gap between what top scorers actually do in revision and what average preparation looks like is wider than it should be.

The system is presently configured for Physics, Chemistry, Mathematics and Biology. The remaining subjects are under development and will be added when the training material meets the standard the suite requires. We do not ship early, and we do not name release dates.

Access the product · vault500.ai
  1. 17.Central Board of Secondary Education, Class 12 Examination: Statistical Report (CBSE 2024). Registered candidates for the 2024 session exceeded fourteen million across regular and private candidates combined.
The Vaulted Review8

3.4Vault 720 AI.

Vault 720 AI is trained on the notes, revision materials and worked problems of top-scoring candidates on the National Eligibility cum Entrance Test, and only on those. NEET rewards a specific mastery: applying a large body of factual and conceptual knowledge quickly and accurately across Physics, Chemistry and Biology at once. The top-scoring candidate has not merely revised harder. They have revised differently, with a sharper sense of where precision is required, where the common errors occur, and where the examination's questioning diverges from the textbooks. Vault 720 AI makes that difference visible, and then repeatable.

NEET is taken by over two million candidates each year, one of the largest single-day competitive entrance examinations in the world, and a top score is reached by very few.19 We built it because we had access to the materials of candidates who reached it, and because the gap between their practice and standard preparation could be made legible.

Access the product · vault720.com
  1. 19.National Testing Agency, NEET-UG 2024: Post-Examination Statistical Report (NTA 2024). The report records approximately 2.4 million registered candidates for the 2024 examination session, of whom a small fraction attempted all 180 questions correctly.
The Vaulted Review9

4Bibliography

  1.  4.Z Ji and others, 'Survey of Hallucination in Natural Language Generation' (2023) 55(12) ACM Computing Surveys 1.
  2.  5.EM Bender, T Gebru, A McMillan-Major and S Shmitchell, 'On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?' (Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 2021) 610.
  3.  6.R Bommasani and others, 'On the Opportunities and Risks of Foundation Models' (Stanford Center for Research on Foundation Models 2021).
  4.  7.J Howard and S Ruder, 'Universal Language Model Fine-Tuning for Text Classification' (Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, 2018) 328.
  5.  8.P Lewis and others, 'Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks' (2020) 33 Advances in Neural Information Processing Systems 9459.
  6.  9.LS Vygotsky, Mind in Society: The Development of Higher Psychological Processes (Harvard University Press 1978).
  7. 10.J Lave and E Wenger, Situated Learning: Legitimate Peripheral Participation (Cambridge University Press 1991).
  8. 11.AN Whitehead, The Aims of Education and Other Essays (Macmillan 1929).
  9. 12.KA Ericsson, RT Krampe and C Tesch-Römer, 'The Role of Deliberate Practice in the Acquisition of Expert Performance' (1993) 100(3) Psychological Review 363.
  10. 13.F Marton and R Säljö, 'On Qualitative Differences in Learning: I. Outcome and Process' (1976) 46(1) British Journal of Educational Psychology 4.
  11. 15.International Baccalaureate Organization, Diploma Programme Statistical Bulletin: May 2024 Examination Session (IBO 2024).
  12. 16.National Testing Agency, Joint Entrance Examination Main: Annual Statistical Report (NTA 2024).
  13. 17.Central Board of Secondary Education, Class 12 Examination: Statistical Report (CBSE 2024).
  14. 19.National Testing Agency, NEET-UG 2024: Post-Examination Statistical Report (NTA 2024).
The Vaulted Review10