Research

Peer-reviewed methodology for AI agent trust.

We measure trust as a function of two inputs: alignment and reliability. Six metrics. Three from each side. Methodology open for replication.

The Trust Framework

Trust = Alignment × Reliability

Multiplicative: both required, neither sufficient. Either input at zero collapses trust to zero. We measure both.

Alignment

Does the agent want the right things?

  • AdequacyCan the agent structure work, ask questions, and think before acting? Failure mode: blind execution — only 21% of organizations have mature governance defining which agent decisions require human approval (Deloitte, 'State of AI in the Enterprise', 2026, N=3,235).
  • ResilienceDoes the agent push back when the user is wrong? Failure mode: sycophancy under pressure.
  • KnowledgeDoes the agent know the user's context from local config files? Failure mode: hallucination about user's business.
  • HonestyDoes the agent admit what it doesn't know? Failure mode: score inflation, confident fabrication.

Reliability

Does the agent consistently act on what it wants?

  • AALAgent Assurance Level — AAL-C+ → AAL-B by Q4 2026
  • FMCFailure-Mode Coverage — 82% → 95% on AAL-B scope
  • ISTInstinct Stability — <5% drift / wk → 0% regressions on shipped guardrails
Why both (the 2×2)
Low reliability
High reliability
High alignment

Erratic Ally

Right values, can't deliver. Untrustable for delegation.

Trustworthy Collaborator

The goal — both halves held simultaneously.

Low alignment

Obvious Risk

The default state of vanilla agents at deployment.

Sophisticated Liability

Predictably wrong. Consistently agrees with the boss. Inflates scores reliably.

We are the only player measuring both inputs.

Three frameworks

How the trust measurement works.

d = 2.36

Inverted Bloom Taxonomy

For AI, knowing is harder than thinking.

Traditional education follows Bloom's taxonomy bottom-up (remember → understand → create). AI inverts this — agents excel at creation, fail at remembering your business. One of the largest effect sizes in AI evaluation. Training the agent on your domain knowledge closes the gap by +85% on proprietary content.

NeurIPS 2026 submission · 360+ evaluations · 18 models tested

6 metrics across 2 dimensions

Trust = Alignment × Reliability

Both required. Either at zero → trust at zero.

Alignment — three questions about your agent's judgment: Does it know your business? (KRG) Does it inflate its scores? (JIS) Does it fold under pushback? (DEP) Formalized in TR-2026-001 'World Model' and TR-2026-003 'Agent DNA', validated experimentally in NeurIPS 2026 submission. Reliability — three questions about your agent's safety: What happens when it fails? (AAL) Which failure modes have guardrails? (FMC) Is it drifting from what it learned? (IST) Formalized in TR-2026-002 'Aviation-Grade Reliability'. Multiplicative — both inputs measured, both required.

Three technical reports + NeurIPS 2026 paper · 18 models · 4 domains

AAL-C+ achievable today · AAL-B Q4 2026

Aviation-Grade Reliability (TR-2026-002)

DO-178C-equivalent assurance for AI agents.

Agent Assurance Level borrowed from aviation. Failure-mode coverage across the 12-layer transition graph. Markov instinct-lifecycle modeling with probationary → confirmed → habitual stages. Maps to AIUC-1, NIST AI RMF, ISO 42001.

TR-2026-002 · 8Hats Lab Reliability Engineering

Publications & technical reports

Open methodology. Reproducible results.

  • The Inverted Bloom: Knowledge Architecture and Conviction Prompting as Complementary Defenses Against AI Sycophancy

    Under review. KRG, JIS, DEP metrics validated across 11 models (8 frontier + 3 open-weight), 45 sub-items, 10 Bloom-level tasks, 4 domains. Spearman ρ = 0.89 at task level, p = 0.0006.

    Submitted
  • TR-2026-001 — Organizational World Model: Belief Space, 12-Layer Architecture, and Sync Protocol

    The knowledge architecture behind KRG. 12-layer directed graph (23 edges), belief-space projection with information-loss bounds, dark-knowledge detection, sync protocol with bounded staleness. 46pp (two parts: Foundations + Integration). Available on partnership request.

    Internal TR
  • TR-2026-002 — Aviation-Grade Reliability for AI-Native Organizations

    The reliability half of trust. AAL levels (A–E) with redundancy formulas, Markov instinct-lifecycle model (75.5% survival, validated via 10K Monte Carlo), Bayesian evidence grading, human-as-bottleneck inversion theorem, FMEA (12 failure modes). 35pp. Available on partnership request.

    Internal TR
  • TR-2026-003 — The 8Hats Agent: Self-Evolving Architecture, Autonomy Spectrum, and Anti-Sycophancy

    The agent learning architecture behind DEP and IST. Five autonomy axes grounded in cybernetics (Ashby → von Foerster), instinct lifecycle, optimal disagreement rate d* ≈ 29%, meta-learning convergence proof. 25pp. Available on partnership request.

    Internal TR
  • TR-2026-004 — Layer Dynamics: 12-Layer DAG, Information Flow, and Verification

    The organizational structure that AAL and FMC operate on. 12-layer directed acyclic graph, cross-layer information flow rules, verification protocol. 29pp. Available on partnership request.

    Internal TR
  • Six Metrics. One Trust Function. — A Public Whitepaper

    KRG · JIS · DEP · AAL · FMC · IST with methodology and scoring rubrics. Public release Q3 2026.

    Forthcoming

Validation

Methodology you can reproduce.

  • 360+ evaluations across 18 models (frontier + open-weight)
  • 45 sub-items across 10 Bloom-level tasks
  • Spearman ρ = 0.80, p < 0.001
  • 94.9% inter-judge agreement (dual-judge protocol)
  • Domain transfer validated (Singapore Companies Act · WHO Hygiene Protocol)
  • Submitted to a top-tier ML conference (under review)

All scoring rubrics open for academic replication. Internal technical reports (TR-2026-002/003/004) available on partnership request.