Trust = Alignment × Reliability. Industry built one half.

A research lab measuring the missing half — three peer-reviewed alignment metrics across 13 benchmarked models.

NeurIPS 2026 · 94.9% inter-judge agreement — Evaluations & Datasets Track, in review

Read the research →Schedule a call

Research — alignment metrics

We measured the missing half.

Three peer-reviewed metrics make agent judgment auditable. Numbers below are from the methodology submitted to NeurIPS 2026.

KRG · Knowledge-Reasoning Gradient

d = 0.04 → 2.36

How much grounding outperforms a vanilla model as task knowledge-dependency rises.

Spearman ρ=0.89, p=0.0006. 11 models, 4 domains.

JIS · Judgment Inflation Score

1.83 ± 0.42

Ratio of model score to ground truth — how much your AI flatters you.

JIS=1.0 calibrated. JIS=2.6 inverts decisions. HALA grounded: 1.00.

DEP · Disconfirming Evidence Prompt

−48 pp

Pressure-sycophancy reduction from a single-sentence intervention (Claude Sonnet 4).

65–87% pressure cut on RLHF models. Zero effect on factual sycophancy.

Validated across 13 models, 4 domains, 3 benchmarks. Inter-judge agreement 94.9%. NeurIPS 2026 Evaluations & Datasets Track (in review).

Read the research →

The position

Labs make agents smarter. Governance makes agents safer. We make organizations ready.

Co-author the trust standard →Email the lab