M-Class Harness
A governed evaluation methodology, scoring framework, and review standard for coding, reasoning, and multi-step AI workflows.
As AI agents move beyond single responses, evaluation must measure sustained work: state tracking, tool use, failure recovery, handoff quality, and evidence trails — under a consistent, governed scoring standard.
Problem Space
Most AI evaluations are too short to expose failures in sustained reasoning, context retention, tool discipline, and recovery from bad intermediate states — and lack a consistent review standard for scoring what they do capture.
System Direction
M-Class defines the evaluation methodology and scoring framework: scored artifacts, evidence-led review patterns, and governed review standards. Execution environments for long-horizon tasks are provided separately by the Long-Horizon Harness; runtime supervision belongs to TripSitter.
Public Capabilities
- 01Evaluation methodology and review standards
- 02Scoring frameworks for sustained work
- 03Coding and reasoning task review
- 04Evidence-led artifact inspection
- 05Public-safe benchmark packaging
M-Class is presented publicly as an evaluation research program. Internal scoring rubrics, prompts, traces, and harness mechanics are not disclosed.
What Is Not Disclosed
Private implementation details, security-sensitive internals, and unreleased runtime architecture are intentionally not disclosed.