Posterior Twins
Distributional Behavioral Simulation for Enterprise Decisions
Enterprise behavioral simulation goes beyond response generation. A useful simulator reproduces the shape of a population under a decision: who accepts, who defects, who hesitates, which segments move, and how much uncertainty remains.
By Ankit Das - Founder & CTO, Twinning Labs

| System | Role | Modal Accuracy | W1 |
|---|---|---|---|
| TL-Twin Alpha | Population-movement model | 0.705 | 1.16 |
| TL-Twin Beta | Calibrated behavior model | 0.728 | 2.20 |
| TL-Twin Gamma | Ensemble scenario model | 0.749 | 2.28 |
| TL-Twin Delta | Direct scenario model | 0.750 | 2.36 |
| Claude Opus 4.7 | Frontier model | 0.770 | 2.87 |
| GPT-5.5 | Frontier model | 0.762 | 3.99 |
| Claude Opus 4.6 | Frontier model | 0.761 | 4.18 |
| GPT-5.4 | Frontier model | 0.749 | 2.59 |
| Gemini 3.1 Pro | Frontier model | 0.742 | 4.19 |
Modal accuracy and distributional fidelity measure different capabilities
A Posterior Twin is a memory-grounded digital twin represented as an updated distribution over likely behavior under a specific decision context. It does not answer as a generic persona; it uses governed customer evidence to estimate what a customer, account, or segment is likely to choose, say, or do, and how uncertain that estimate is.
The public result set uses the comparable 226-example held-out evaluation. TL-Twin Alpha achieves the lowest observed Wasserstein-1 distance, while TL-Twin Delta and TL-Twin Gamma provide balanced operating points near the modal-accuracy frontier.
Enterprise decisions live in distributions, not isolated responses
Synthetic audiences, AI-moderated research, digital-twin systems, multi-agent population simulators, and customer-data decisioning products all approach the same enterprise problem from different surfaces.
The technical boundary that matters is whether the system gives the decision-maker defensible response distributions under a specified decision context: decision direction, behavioral generation, population shape, and decision trace.
Posterior Twins combine memory, model operating points, and simulation
The TL model family supplies the operating frontier measured in this paper. The Memory Layer connects governed enterprise evidence across customer, product, research, commercial, and outcome systems, then preserves that evidence as stable twin context.
The Simulation Engine creates scenario environments around those twins, maps outcomes to the scenario contract, aggregates distributions or structured generative artifacts, and exposes the result as an auditable decision object.
Two metrics answer two different enterprise questions
Modal accuracy compares the simulator with the empirical human mode: the behavior selected most often by people. This matters when teams need decision direction, such as which message, offer, package, or product option wins on average.
Wasserstein-1 distance asks how far the simulated population curve must move to match the empirical human curve. This matters when allocation, pricing, risk, and launch decisions depend on the population underneath the headline direction.
The benchmark is an operating frontier, not a single leaderboard
Claude Opus 4.7 has the highest frontier-model modal-accuracy point estimate in the public result set. GPT-5.4 has the strongest frontier-model W1 point estimate. TL-Twin Alpha has the lowest observed W1 overall.
TL-Twin Gamma matches GPT-5.4 at the same rounded modal accuracy while improving W1 from 2.59 to 2.28. TL-Twin Delta reaches 0.750 modal accuracy with W1 = 2.36.
The system turns the frontier into reusable decision infrastructure
Posterior Twins need four assets working together: governed memory, memory-grounded digital twins, a simulation engine, and a decision record that can be reused after the run.
The returned artifact contains the routed operating point, result distribution or generated trace, scenario assumptions, and enough context for review. This lets teams compare scenarios, attach evidence to a decision record, and revisit the run when observed outcomes arrive.
Behavioral simulation should be measured, routed, and governed
The defensible asset is not the ability to prompt a frontier model into sounding like a customer. The defensible asset is governed memory, a calibrated behavioral model family, a repeatable simulation engine, and measurement discipline.
Posterior Twins move enterprise AI from generating what a plausible respondent might say to simulating how a governed population is likely to move under a decision.

Das, A. (2026). Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions. Twinning Labs Research, v1.
@techreport{das2026posterior,
author = {Das, Ankit},
title = {Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions},
institution = {Twinning Labs},
year = {2026},
version = {v1}
}Modal accuracy is not distributional fidelity
Read the journal essay that frames the paper for operators, investors, and technical buyers.