Methodology + strongest limitations

Match before comparing

Each record keeps its date, prompt, conditions, provenance, quotation boundary, source, page locator, coding scope, and evaluator role attached. Differences are compared only where controls permit.

The unit

The research unit is a retained short excerpt attached to whole-response provenance. Historical and 1979 codes apply to the quoted excerpt. Current codes apply to full agency-published responses, although this site reproduces only a short consecutive passage. Scope travels with every record.

Honest periodization

TargetEvidence actually usedRule
Near 19261919 Virginia core; 1928 published supplements of unknown writing yearNever “1926 essays”
Near 1976Spring 1979 NAEP age-17 responsesNever “1976 homework”
PresentOperational 2024–2025 state anchorsNever “2026 samples”

Selection hierarchy

  1. Calibrated core: five Hudelson scale points around a Virginia survey median and ten interior 1979 NAEP task exemplars.
  2. Operational grade-level anchors: six Florida 3/3/3 papers explicitly within grade-level range.
  3. Robustness anchors: four Massachusetts midpoint-ideas/top-conventions and two Texas partial-development/top-conventions papers.
  4. Average-rated supplements: four 1928 seventh/eighth-grade textbook pieces with only an aggregate “Fair plus or Good” label.
  5. Uncalibrated supplements: five textbook-selected pieces with unknown writing dates, prompts, scores, and editorial histories.

The weaker 1928 material is never collapsed into the calibrated 1919 core. Every displayed anchor is purposively published, even when its underlying assessment frame is broad.

Comparison panels

A. Timed, source-free narrative

Five 1919 narratives and two 1979 picture narratives form the closest performance match. Age, geography, prompt type, archive medium, score construct, and selection still differ. There is no retained current match.

B. Persuasion under unlike evidence conditions

Two brief 1979 source-free speeches are read beside current source-based arguments. This panel describes what each task elicits and rewards; it cannot estimate a gain.

C. Present source-based analysis

This is a descriptive present-day lane, not a trend. Missing historical counterparts do not prove historical students never did such work.

D. Unmatched forms

Practical, humorous, sensory, evaluative, literary, and informational forms identify breadth but stay outside matched claims.

Coding discipline

Codes describe visible conceptual claims, idea relationships, abstraction, causal reasoning, qualification, evidence use, organization, sentence control, lexical precision, and articulation limits. They never code intelligence, character, diligence, demographic capability, or unseen causes. Categories are descriptive—not equal-interval scores—and no composite is allowed.

Observation: directly visible language or a documented condition. Inference: a bounded interpretation tied to those observations. Speculation: an untested causal story, kept out of findings.

Strongest limitations

What these limits rule out

No national average, rise/decline, statistical significance, annualized rate, error-rate series, intelligence claim, demographic judgment, teaching-quality judgment, or cultural cause can be recovered from this design.

Reproduce the audit

The authoritative record remains in the repository’s three data/samples-*.json files and data/context.json. Generated CSV and summaries are checked for freshness. The source register preserves source IDs, direct URLs, locators, and access dates. Start with the browsable source library, then inspect the repository for scripts and full research notes.