Skip to main content
Twinning Labs
PlatformResearchJournalContactLoginBook Demo

Posterior Twins

Distributional Behavioral Simulation for Enterprise Decisions

Technical Whitepaperv1May 7, 202630 min

Enterprise behavioral simulation goes beyond response generation. A useful simulator reproduces the shape of a population under a decision: who accepts, who defects, who hesitates, which segments move, and how much uncertainty remains.

By Ankit Das - Founder & CTO, Twinning Labs

Operating frontier comparing modal accuracy and distributional fidelity
Operating frontier on the held-out evaluation. Lower W1 is better; higher modal accuracy is better.
Public Result Set
SystemRoleModal AccuracyW1
TL-Twin AlphaPopulation-movement model0.7051.16
TL-Twin BetaCalibrated behavior model0.7282.20
TL-Twin GammaEnsemble scenario model0.7492.28
TL-Twin DeltaDirect scenario model0.7502.36
Claude Opus 4.7Frontier model0.7702.87
GPT-5.5Frontier model0.7623.99
Claude Opus 4.6Frontier model0.7614.18
GPT-5.4Frontier model0.7492.59
Gemini 3.1 ProFrontier model0.7424.19
Executive Summary

Modal accuracy and distributional fidelity measure different capabilities

A Posterior Twin is a memory-grounded digital twin represented as an updated distribution over likely behavior under a specific decision context. It does not answer as a generic persona; it uses governed customer evidence to estimate what a customer, account, or segment is likely to choose, say, or do, and how uncertain that estimate is.

The public result set uses the comparable 226-example held-out evaluation. TL-Twin Alpha achieves the lowest observed Wasserstein-1 distance, while TL-Twin Delta and TL-Twin Gamma provide balanced operating points near the modal-accuracy frontier.

Category Problem

Enterprise decisions live in distributions, not isolated responses

Synthetic audiences, AI-moderated research, digital-twin systems, multi-agent population simulators, and customer-data decisioning products all approach the same enterprise problem from different surfaces.

The technical boundary that matters is whether the system gives the decision-maker defensible response distributions under a specified decision context: decision direction, behavioral generation, population shape, and decision trace.

Approach

Posterior Twins combine memory, model operating points, and simulation

The TL model family supplies the operating frontier measured in this paper. The Memory Layer connects governed enterprise evidence across customer, product, research, commercial, and outcome systems, then preserves that evidence as stable twin context.

The Simulation Engine creates scenario environments around those twins, maps outcomes to the scenario contract, aggregates distributions or structured generative artifacts, and exposes the result as an auditable decision object.

Evaluation Method

Two metrics answer two different enterprise questions

Modal accuracy compares the simulator with the empirical human mode: the behavior selected most often by people. This matters when teams need decision direction, such as which message, offer, package, or product option wins on average.

Wasserstein-1 distance asks how far the simulated population curve must move to match the empirical human curve. This matters when allocation, pricing, risk, and launch decisions depend on the population underneath the headline direction.

Results

The benchmark is an operating frontier, not a single leaderboard

Claude Opus 4.7 has the highest frontier-model modal-accuracy point estimate in the public result set. GPT-5.4 has the strongest frontier-model W1 point estimate. TL-Twin Alpha has the lowest observed W1 overall.

TL-Twin Gamma matches GPT-5.4 at the same rounded modal accuracy while improving W1 from 2.59 to 2.28. TL-Twin Delta reaches 0.750 modal accuracy with W1 = 2.36.

System

The system turns the frontier into reusable decision infrastructure

Posterior Twins need four assets working together: governed memory, memory-grounded digital twins, a simulation engine, and a decision record that can be reused after the run.

The returned artifact contains the routed operating point, result distribution or generated trace, scenario assumptions, and enough context for review. This lets teams compare scenarios, attach evidence to a decision record, and revisit the run when observed outcomes arrive.

Conclusion

Behavioral simulation should be measured, routed, and governed

The defensible asset is not the ability to prompt a frontier model into sounding like a customer. The defensible asset is governed memory, a calibrated behavioral model family, a repeatable simulation engine, and measurement discipline.

Posterior Twins move enterprise AI from generating what a plausible respondent might say to simulating how a governed population is likely to move under a decision.

Distributional-fidelity scorecard for measured operating points
Distributional-fidelity scorecard for measured operating points, with modal accuracy kept visible under each bar.
Suggested Citation

Das, A. (2026). Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions. Twinning Labs Research, v1.

@techreport{das2026posterior,
  author = {Das, Ankit},
  title = {Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions},
  institution = {Twinning Labs},
  year = {2026},
  version = {v1}
}
Companion Essay

Modal accuracy is not distributional fidelity

Read the journal essay that frames the paper for operators, investors, and technical buyers.

Read Essay