Skip to main content
Twinning Labs
PlatformResearchJournalContactLoginBook Demo
Back to Journal

Posterior Twins: Modal Accuracy Is Not Distributional Fidelity

May 7, 20263 min readResearch
Read the published whitepaper
Posterior Twins: Modal Accuracy Is Not Distributional Fidelity

TL-Twin Alpha hits the lowest Wasserstein-1 in the public result set: W1 = 1.16. No frontier model gets close.

Enterprise pricing, product, and marketing decisions live in the shape of a population, not the median answer. Which segments delay, downgrade, expand, defect? Whose enthusiasm comes at the cost of whose alienation? Modal accuracy cannot answer those. Distributional fidelity can.

That is what we built Posterior Twins to deliver, and the numbers back it up.

The 226-example evaluation, headline points:

  • TL-Twin Alpha: lowest observed W1 of 1.16. Twinning Labs population-movement model. Frontier-leading.
  • TL-Twin Delta: modal accuracy 0.750, W1 2.36. Direct scenario model.
  • TL-Twin Gamma: modal accuracy 0.749, W1 2.28. Ensemble scenario model.
  • Claude Opus 4.7: best frontier-model modal accuracy at 0.770, W1 = 2.87.
  • GPT-5.4: best frontier-model W1 at 2.59 (still 2.2x Alpha).

No single frontier model dominates every metric. The TL-Twin family does, by routing.

Why W1 matters more than the leaderboard. Wasserstein-1 measures allocation error across the population curve. A model can call the leading offer correctly and still misplace demand, overstate enthusiasm, or miss the rejection tail. Budgets, pricing, sales capacity, and risk exposure live in that tail.

Posterior Twins is not a persona prompt. It is an updated distribution over likely behavior under a decision context. Three parts:

  1. Memory Layer: governed enterprise evidence connected across customer, product, commercial, research, and outcome systems, preserved as durable twin context.
  2. Simulation Engine: turns memory state into repeatable scenario runs and measured distributions.
  3. TL-Twin model family: Alpha (population movement), Beta (calibrated behavior), Gamma (ensemble scenarios), Delta (direct scenarios). The system routes; teams do not pick models manually.

A generic model call starts over each time. A Posterior Twin compounds: memory, scenarios, traces, and measured distributions become reusable decision infrastructure. That is the difference between a synthetic-response toy and a behavioral-simulation system mature enterprises can deploy in their own cloud.

Read the full whitepaper for the methodology, the held-out protocol, and the per-segment results: twinninglabs.com/research/posterior-twins.

If your decisions depend on the shape of a population, reach us at [email protected].

- Ankit Das, Founder & CTO, Twinning Labs

Share this story

If this insight was useful, share it with your team.