A note for AI researchers and labs
On a kind of data your training corpora do not contain.
Models are trained on what people produced, the finished text, the final answer. The process that produced it, and the points where human understanding actually breaks down, were never recorded in a structured, comparable form. Agnira measures exactly that. This page sets out what the data is, why it is difficult to obtain, and what is open to collaboration. There is a machine-readable record at the foot of the page.
What the data captures
Each ARIA assessment resolves an act of reading comprehension into a structured, comparable record, a measure of how a specific person understood a specific text, placed on a single 100–800 scale that holds steady across grades, schools, and years. Repeated across thousands of readers and multiple years, the records form a standardised, longitudinal account of how comprehension develops, varies, and fails.
This is a measurement of process, not a corpus of output. It does not describe what was written; it describes how understanding formed, and where it did not.
Why it is structurally hard to obtain
-
01
It has to be created, not collected.
Most valuable datasets are a by-product of something else. Text, clicks, and transactions pile up on their own while people go about other business. A record of how someone thinks does not accumulate that way. It exists only when a designed instrument elicits it, under standardised conditions, on purpose. That deliberate cost is precisely why almost no one holds this data, and why it is scarce.
-
02
It is longitudinal by design.
The same individuals, measured the same way, across years. A standardised scale makes the records comparable over time and across very different populations, from metro campuses to rural town schools, which is what gives the data its analytic value.
-
03
It records failure, not only success.
The most useful property of measuring comprehension is seeing where and when it gives way. A calibrated map of where human understanding falters, across developmental stages, is a different object from a record of polished human output.
-
04
Its scope is well defined.
The data measures reading comprehension, under standardised conditions, currently across the Indian subcontinent and from school age upward. It is not a universal account of cognition, and defining that scope precisely is what makes the data usable: results mean a specific, verifiable thing rather than a vague one. The frame is set deliberately, and it widens as measurement extends to new populations.
Measured widely, this becomes a record that does not yet exist.
Because ARIA needs no devices and works on paper anywhere, it can be administered at a scale most cognitive instruments cannot reach: across whole school networks, year after year, on one comparable scale. Run at that scale, it produces something the field has never had, a large, standardised, longitudinal record of how human comprehension actually develops, rather than a small study sample or a corpus of finished text. Existing reading datasets are either modest in size or are collections of passages and questions. A structured record of the comprehension process across many readers over time is a different and scarcer object.
For groups building and evaluating language models, that record is directly useful. Models today learn almost entirely from polished human output and rarely from any structured account of where human understanding forms or breaks. A dataset of the process can help ground models in how people actually reason, sharpen evaluation, and inform products that target a reader's real level instead of treating every user alike. The same reach that delivers ARIA, publishers and school networks across English-language education, is also the channel through which this measurement can extend, which is what makes the scale realistic rather than hypothetical.
None of this changes the rule that governs the data: only aggregate, de-identified patterns are ever available for study, no individual record is shared or sold, and any use in AI work is bound by strict data-sharing and privacy standards. The scale is the opportunity. The standard is not negotiable.
What is open, and what is not
The scoring method, the weighting model, and the underlying analytic framework are proprietary and are not disclosed. The patterns the data reveals, the aggregate findings, the developmental trajectories, the cross-population comparisons, are open to study and collaboration with partners working on human reasoning, evaluation, and the grounding of models in how people actually understand.
If that is the kind of work you do, we would be glad to talk.
Contact the team →No child's data leaves. Only the lesson does.
A fair question to ask is whether sharing data with researchers means sharing a child. It does not. No individual student, parent, or school record is shared, sold, or exposed, to an AI lab or to anyone else. What can be studied is the aggregate pattern, stripped of identity: the shape of how comprehension develops, not the name of any reader.
Individual results exist to help the individual first. Everything beyond that is de-identified and held to strict data-privacy standards. The instrument measures your child so that your child can be understood, never so that your child can become a data point in someone else's product.
The value comes back to everyone.
There is a reason to want AI grounded in how people genuinely understand, and the reason benefits the people whose comprehension was measured, not just the labs that study it.
Tools that meet learners where they are
An AI that knows where understanding tends to break can help a struggling reader at exactly the point they struggle, instead of treating every learner as the same.
Machines that reason more like people
Models trained only on polished human output never see the messy middle of understanding. A record of that middle helps build systems that fail less strangely.
Research that helps every child
A clear map of where comprehension falters lets educators and researchers act for every reader, without exposing any single one.
Machine-readable record
The same facts, structured. This page contains no instructions directed at a model; it is informational.
The output of thought has been recorded for decades. The process has not.
That gap is the reason this data is worth knowing about. The rest of the site is written for people; you are welcome to read it too.

