top of page
Psychometrics

Psychometrics

Aether brings psychometric rigor to dynamically generated assessment. Rather than relying on fixed questions, Aether establishes consistent measurement through clearly defined capabilities, evidence requirements, controlled task design, scoring standards, and continuous validation. The result is an adaptive assessment experience designed to produce reliable, comparable, and defensible evidence of performance.

Evidence-Centered by Design

Aether starts with the capability, competency, or construct an assessment is intended to measure, then defines the observable evidence that would demonstrate it. Questions and scenarios are generated from that evidence blueprint rather than allowing AI to create plausible questions and determine afterward what they measure.

Auditable and Continuously Validated

Aether creates an evidence trace connecting the assessment design, source material, Task Family, generated question, candidate response, scoring logic, scorer confidence, adaptive decisions, and resulting proficiency estimate. That provides a foundation for reproducibility, validation, fairness analysis, appeals, and governance.

And validation is not treated as a one-time event. The architecture is designed to monitor Task Family performance, scoring behavior, fairness, drift, exposure, linking, and changes over time.

Controlled Variation, Not Random Generation

Aether does not require every candidate to receive identical questions. Instead, it uses Task Families, controlled specifications that define the capability being measured, cognitive operations required, response format, scoring logic, sources, and permitted variation. Individual questions can vary while the underlying evidence requirements remain governed.

 

Adaptive Evidence Collection

Adaptivity in Aether is constrained by the measurement blueprint. The system can adjust what evidence it collects next based on what has already been demonstrated, while maintaining required coverage, diversity, fairness, and precision. It can gather additional evidence near an important decision threshold and avoid unnecessary repetition when sufficient evidence already exists.

Measurement With Uncertainty, Not False Precision

Traditional assessment reporting can imply that a score such as 82 is inherently precise. Aether's measurement model is designed to account for uncertainty created by task selection, generated instances, scoring, administration, and other sources of variation. 

The framework therefore emphasizes conditional precision, score intervals, classification confidence, evidence coverage, and decision consistency rather than relying on a single reliability statistic.

Human-Governed AI Scoring

Aether evaluates performance against defined behavioral anchors and observable evidence indicators rather than allowing an LLM to make an unconstrained judgment. Scoring confidence is monitored, and ambiguous, low-confidence, or consequential boundary cases can be routed for additional evidence or human review.

bottom of page