A formal assessment has fixed content, a fixed procedure, published scoring rules, and documented evidence behind its scores. Those four conditions are the whole difference. They are also why the results travel: across classrooms, across schools, and across years.
Below you’ll find the four conditions, the two scoring frames, and what validity and reliability actually mean after that: how to read a score report, and the accommodation rules that apply under IDEA and Section 504.
| Feature | What it means in practice |
| Fixed content | Every student sees the same items, in the same order |
| Standardized administration | Same script, same timing, same materials, same setting rules |
| Published scoring rules | The result does not move depending on who grades it |
| Technical evidence | The manual reports validity and reliability data |
| Score frame | Read against a norm group, or against a fixed standard |
| Who gives it | Staff trained on that specific instrument |
| What it decides | Eligibility, placement, accountability reporting, multi-year progress |
Direct answer: A test counts as formal when the publisher fixes its content, administration, and scoring, and the manual reports validity and reliability evidence. You then read the score against a norm group or a fixed standard, which is what lets schools use it for eligibility, placement, and accountability decisions.
Key takeaways
- Four conditions make a test formal, and length is not one of them.
- Norm-referenced and criterion-referenced scores answer different questions. Neither answers both.
- Reliability asks if the number would hold next week. Validity asks if it supports your decision.
- Every score report carries a confidence band, because every score carries measurement error.
- An accommodation preserves the score. A modification invalidates it.
- Federal rules put initial evaluations on a 660-dayclock and forbid deciding eligibility on one measure.
The Four Conditions That Make a Test Formal

Length and difficulty have nothing to do with it. A ninety-minute final exam written by one teacher is still informal in the technical sense, and a six-minute reading screener can be fully formal. What matters is whether the instrument is fixed and documented.
- Fixed content. Every student sees the same items, in the same order, at the same difficulty. Nobody swaps a question because a class struggled with it last period.
- Standardized administration. The manual sets the script, timing, materials, seating, and what an examiner may say when a student asks for help.
- Published scoring rules. Rubrics, answer keys, and conversion tables come from the publisher, so two scorers reach the same result.
- Technical evidence. The manual reports validity and reliability data, plus either a norm sample or a defined performance standard.
Drop any one of the four and comparability collapses. That loss is practical, not bureaucratic.
Norm-Referenced and Criterion-Referenced Scores Answer Different Questions
This distinction causes more confusion in parent meetings than anything else on the page. One table fixes it.
| Norm-referenced | Criterion-referenced | |
| Question it answers | How does this student compare with peers? | Has this student met a defined standard? |
| Score you see | Percentile rank, standard score, stanine | Proficient, meets standard, 82 percent correct |
| Built from | A national sample tested under identical conditions | A blueprint of skills tied to the curriculum |
| Good for | Eligibility, spotting an outlier, tracking rank over years | Grading, course placement, checking mastery |
| Cannot tell you | Which specific skill is missing | Whether the student is ahead of or behind peers |
| Typical example | Cognitive and achievement batteries used in evaluations | State proficiency tests, end-of-course exams, licensure exams |
Say a report shows a percentile rank of 30. That student scored higher than 30 percent of the norm group. It does not mean 30 percent correct, and it names no skill to teach. When a result lands well above the norm group, the useful reply is a harder task, not a louder celebration. That is where deeper math enrichment ideas come in.
Validity and Reliability, Translated
Both words sit in every test manual and rarely in plain English.
Reliability asks whether the same student would get roughly the same score next Tuesday. Validity asks whether the score supports the decision you want to make with it. A test can be highly reliable and still invalid for your purpose. Spelling tests measure spelling consistently, and tell you little about whether a child should skip a grade.
So turn the jargon into two questions before any meeting. Would this number hold up on a different day? Does it actually speak to the decision on the table? If either answer is no, ask for a second source of evidence.
Standardized Administration Is the Part Schools Get Wrong
Scores stop being comparable the moment conditions drift. Every formal assessment ships with a manual, and that manual is the rule book. Drift usually looks like one of these:
- Reading a math item aloud when the rules forbid it.
- Granting extra time nobody approved.
- Prompting a student after a wrong answer.
- Testing a child in a hallway because no room was free.
Remote settings add their own problems: an unproctored device, a sibling in the room, a slow connection eating the clock. Families weighing full-time virtual school programs should ask early how the school proctors its required tests, because the answer varies enormously.
Trained examiners matter for the same reason. Someone who has not practiced the basal and ceiling rules on a given battery will produce a number. It just will not be a trustworthy one.
How to Read a Score Report Without Overreading It

Most reports show four things: a raw score, a standard or scaled score, a percentile rank, and a confidence band.
The band is the honest part. Every score carries a standard error of measurement, so the report prints a range instead of pretending the number is exact. Take a standard score of 92 printed with a band of 87 to 97. The student’s true score most likely sits somewhere in that stretch. Two students five points apart may not differ at all.
Grade equivalents are the trap. Your fourth grader with a grade equivalent of 7.2 has not mastered seventh grade work. She scored roughly what an average seventh grader would score on fourth grade material. Ignore the number.
Aggregate results follow a similar rule. Once a state pools individual results into a building average, ranking sites publish them and fold them into the numbers behind school ratings. One digit then stands in for years of variation.
Accommodations and Modifications Are Not the Same Thing
Both appear in an IEP or a Section 504 plan, and mixing them up can void a score.
| Accommodation | Modification | |
| What changes | How the student reaches the test | What the test measures |
| Examples | Extended time, large print, a scribe, a separate room, a screen reader | Fewer items, easier passages, a word bank, dropped sections |
| Effect on the score | Stays comparable and reportable | No longer comparable to the norm or standard |
| Who approves it | The IEP or 504 team, written into the plan | The same team, and the report has to flag it |
Read-aloud is the one that trips people. Reading a science passage aloud is usually an accommodation. Reading a decoding test aloud strips out the skill under measurement, so it becomes a modification and invalidates the result. Both an IEP and a 504 plan can carry accommodations, though only an IEP brings specialized instruction with it.
When a Formal Assessment Is Legally Required

Teacher judgment carries plenty of weight, right up until a decision has legal consequences. Eligibility for special education is the clearest case.
Federal rules put that decision on a clock. Under the U.S. Department of Education’s IDEA regulations on evaluation timelines, a school must conduct an initial evaluation within 60 days of parental consent. A state may set its own shorter timeline instead. Those decisions affect many children. The National Center for Education Statistics reports that 7.5 million students, 15 percent of public school enrollment, received services under IDEA in the 2022-23 school year.
The U.S. Department of Education’s evaluation procedures for special education add three more rules. Schools may use a test only for the purposes for which it is valid and reliable. Trained and knowledgeable personnel have to administer it. And no single measure may serve as the sole criterion for deciding whether a child has a disability. The same guidance sets reevaluation at least once every 3 years, unless the parent and the school agree it is unnecessary.
Families should quote that last rule most often. Referrals normally start long before any of this, when a teacher notices that targeted instruction is not moving a skill. Six weeks of documented reading fluency practice with no measurable change is exactly the evidence an evaluation team wants in front of it.
Your Next Step
Pull the most recent score report in your file and do three things with it. Find the confidence band and say the range out loud instead of the single number. Check which scoring frame produced the result, then ask what question that frame was built to answer. Write down the one decision the number is being used to justify.
If the number cannot carry that decision by itself, you have found your next question for the team.
If you want to know about Authentic Assessment then visit our Exams category.
Frequently Asked Questions
It is a test with fixed items, a fixed procedure, published scoring rules, and technical data in its manual. Those features let you compare one student’s result against a norm group or a defined standard.
Whoever the publisher’s manual qualifies, that usually means a school psychologist, an educational diagnostician, a speech-language pathologist, or a trained teacher for group-administered tests. Federal rules require trained and knowledgeable personnel for special education evaluations.
Usually not. A teacher-made unit test has no norm sample, no external scoring rules, and no published reliability data, so its results travel no further than that classroom.
Yes. Put the request in writing and date it, because the clock and the school’s duty to respond both start from that date. A district that declines must give written notice explaining why.
Reevaluation under IDEA happens at least every three years, unless the team agrees it is unnecessary. State accountability testing runs annually in the tested grades.
Say so at the meeting, and bring examples. One test session captures one hour of one day, and no team may rest an eligibility decision on a single measure anyway.
