Thoughts. What I notice while building.
"The probability that profit drops by more than twenty percent is 76 percent." That sentence sounds important and is worth almost nothing. The sentence that is worth something goes like this: "76 percent, and 46 percent of that spread disappears if you measure one single assumption you have so far been guessing." Everything in this article sits between those two sentences.
The first sentence is a prediction. Predictions about complex systems are cheap and mostly wrong. I tested a dozen famous economic relationships honestly, out of sample, with confidence intervals. Survivors: none. Anyone still buying a point forecast after that is buying entertainment.
The second sentence is not a prediction. It is a statement about your own ignorance. And that one you can attack, because it tells you where to start.
The number that says nothing
A point forecast is a distribution somebody folded up. "The process step takes 7.5 minutes" is the folded version of "5 to 10 minutes, usually 7". What gets lost in the folding is exactly what you wanted the number for: the range in which you are wrong.
A band instead of a number is the first step, and it is only half of one. Because a band still answers the wrong question. The wrong question is "how will it turn out?" The right one is "how sure are we, and what would we have to know for it to get surer?"
This is not word games. It decides whether the calculation ends in an action or a shrug.
Watch what happens when it unfolds. Scroll slowly.
The strokes are the same before and after. What was missing was never the arithmetic. It was the sentence.
Two ways of not knowing
This is the thesis everything hangs on, and it is uncomfortably simple.
The world rolls dice. Next quarter's demand will be what it will be. No amount of study fixes it. You cannot work this uncertainty away, you can only price it. In the jargon: aleatoric. The lottery piece is about the fiasco that follows from treating this part as predictable.
You do not know. You guessed an effect size. You estimated a p90 instead of measuring it. You took a baseline from a report nobody checked. This uncertainty is removable. Jargon: epistemic. And it is the only part with a bill you can actually pay.
Separating those two is the product. Everything else is trim.
Because in the result they look identical. A wide band can mean "the world is wild" or "I guessed". Same bar on the page, exactly opposite advice. In the first case you build buffers. In the second you pick up the phone and call procurement.
The next drawing is the most important one in the article. Give it the thirty seconds.
Watch the lower bracket. At the end it is exactly as long as it was at the start. Measuring pushed the worlds together and never touched the dice. Confuse the two and you will expect something from data that data cannot deliver.
How to separate them: two loops
The separation is not asserted, it is produced. By nesting.
outer loop, M worlds: draw what you do not know.
effect sizes, priors, baselines, distribution shifts.
inner loop, N draws: draw what the world rolls.
the drivers, correlated, through their distributions.Each world is one complete state of knowledge: "suppose the elasticity really is 0.7 and the baseline holds." Inside that world, the world rolls a thousand times. Then the next world, a different state of knowledge, another thousand draws.
By the law of total variance the spread then falls apart cleanly:
aleatoric = E_world[ Var_inner ] the world rolls; no knowledge fixes it epistemic = Var_world[ E_inner ] you do not know; measuring fixes it
The spread inside a world is the dice. The spread between worlds is your ignorance. A run for me is 512 worlds times 2,000 draws, so a good million simulations, and it takes hundreds of milliseconds. That is why the interface is not a report but an instrument you turn.
A tornado chart cannot do this. It moves one input at a time, deterministically, and it cannot say how much of the band disappears if you measure. It shows you which lever is loud. Not which lever is yours.
The same dice, and what they cost
Every world maps the same random numbers through its own distributions. An inner draw is therefore "the same percentile", not "the same value". Which makes the difference between two worlds purely a difference in knowledge and nothing else.
That is load bearing. The sharpest test of the whole tool is therefore not statistical: comparing a variant against an identical case must produce exactly zero. Not zero in expectation. Zero.
Shared dice are not free. They make the layout crossed, and the standard error falls into three parts instead of two:
SE² = world/M + draw/N + interaction/(M·N)
The middle term is the trap. It is common to all worlds, so no number of worlds shrinks it. Drop it and you understate your own error bar several times over, worst exactly when the tool looks most confident.
So my report does not say "76.1 percent". It says "76.1 percent, give or take 2.5 percentage points". And once that band exceeds five points, the fixed number disappears and it says "too uncertain for a fixed number". A tool whose whole thesis is "know where you are weak" does not get to print false precision about itself.
Three numbers become a distribution
The author of a case gives three numbers: p10, p50, p90. Bad case, middle case, good case. It is the only elicitation format humans are any good at. Nobody can pick a distribution family, and nobody should be asked to.
The family is derived. I use a metalog because it is linear in its coefficients: three quantiles pin three terms through one exact linear solve. No optimiser, no degenerate direction, no starting values.
Why not lognormal or Student-t, the way half the world does it? Both carry a location parameter that chases p10 when you fit three quantiles. The spread explodes to compensate, the middle quantiles still fit nicely, and the tail turns to nonsense. In testing, a log-Student-t for raw material costs with the entirely harmless numbers 0.93 / 1.0 / 1.35 produced a p99.9 of 55 times baseline. That is not a distribution, that is a generator for board meetings.
And then the place where almost every model quietly crashes: the bounds.
A lower bound is not innocent. Under a log transform it does not merely clip the left side, it inflates the right tail. Hold the same three numbers and move the lower bound from 0.0 to 0.9, and the implied p99.9 goes from 2.8 to 13.6.
Nobody expects that. So every fit reports the extreme quantiles it implies, and a bound sitting near p10 raises an advisory.
What does not exist is a "fat tail" dial. A dial nobody can calibrate is worse than a number nobody can hide from.
Drivers fail together, and that is the point
Without correlation, Monte Carlo understates the tails systematically. Chips and staff go missing at the same time. Energy and raw materials rise together. And the tails are the product: nobody builds a simulation to learn the mean.
You declare a rank correlation, because that is the number a human can assert. Then comes the place where most tools go quiet.
The Gaussian copula has zero tail dependence. Under it, the chance that the second driver is also extreme tends to its independent value the further out you go. However high the declared correlation. Joint catastrophes get understated precisely where people read. For a tool whose thesis is the tails, that is a strong claim to make silently.
The t copula fixes it with one extra number. You draw the drivers jointly and they share one additional roll: "how violent is this world". That shared roll is the entire mechanism, because that is what a common shock does. At a correlation of 0.55 and six degrees of freedom, tail dependence is roughly 0.2. Under Gaussian it is 0. Same declared correlation, fundamentally different catastrophe.
This cannot be shown as a scatter. At this correlation both couplings put about the same number of draws in any corner you can still see. The difference only opens up far enough out that a visible cloud has nothing left in it. So here is the quantity that carries it.
On the left the two are indistinguishable. On the right one curve heads down and the other does not. Choose your copula by looking at the middle and you have chosen the tail without looking at it.
What to measure next
Now the part all of this is built for.
For every assumption drawn once per world, you can ask: how much of the spread between worlds would vanish if this one assumption were known? That costs one regression over M points. No extra simulation, no second run, nothing.
The case I developed this on is a carmaker with chip exposure. Five drivers, eight consequences, four folded observations, twelve month horizon. Important before anyone quotes the numbers: this is my model from the outside, with guessed assumptions, and it says nothing about the company. It says something about the model. Which is precisely the point.
The run says:
| Reading | Value |
|---|---|
| P(profit falls more than 20%) | 76.1%, give or take 2.5pp |
| Of that, your own not knowing | 41% |
| Best candidate to measure | demand shock |
| What measuring it removes | about 46%, give or take 6pp |
| Provenance of that driver | a guess with no source |
| What the ranking covers in total | 77% |
Read the second to last line again. The strongest lever in the whole model is a number somebody wrote down without a source for it. That is the normal case, not the exception, and without this calculation you never see it.
Two honesty constraints hang off this, and neither is decoration.
The shares do not sum to one. First order indices ignore interactions. So the sentence is "46 percent of the removable spread would go away if you measure the demand shock alone" and never "46 percent comes from the demand shock". The first is what the estimator supports. The second would be a decomposition claim it cannot make. The ranking also states what it misses: it covers 77 percent, the rest is interaction that no single measurement collects.
No binning. With a few hundred worlds each bin holds few worlds, and the inner sampling noise leaks into the between group variance. The leverage would come out biased upward, and the tool would promise more benefit from measuring than measuring can deliver. The raw shares summed to 91 percent, of which 14 points is what regressions explain from pure noise. Those points are not in the report. They have been subtracted.
And it is not called expected value of partial perfect information. That would be a different quantity and would need a set of decisions, a utility over them and a currency. This tool has none of those, and it does not invent them.
Enough explaining. Here is a smaller case, four drivers, and you get two measurements. The arithmetic underneath is the one described above: worlds on the outside, draws on the inside, split by the law of total variance. It runs in your browser while you click.
A small case with four drivers. The question is the probability that the outcome falls more than twelve percent below baseline. You have budget for two measurements. Each one replaces an estimate with a number. Watch the band, not the headline percentage.
Most people reach for the widest driver. That is the trap, and in real cases it is the rule rather than the exception: wide usually means the world rolling dice. Provenance tells you more than range does. "A guess, no source" is not a blemish on the data sheet. It is the address of your biggest lever.
Ruin is not a bad percentile
A case that says something about time runs over a step grid. Stocks integrate, delayed edges land later, and thresholds can absorb. Absorbing means: once a path crosses the line, it stays across. The cash is gone, the stocks freeze, there is no continuing.
That is not a detail, it is a different calculation. A bad percentile is a value. Ruin is a state.
From which follows the number a single period run fundamentally cannot give you: the gap between two means. The unconditional one counts every path, the ruined ones at the value they carried when they locked. The conditional one counts only those still standing. The gap between them is what ruin costs.
And the chart carries the ruined paths along. Dropping them would let the fan drift upward while the ruined ones quietly vanish from the sample. That would be a chart claiming the opposite of what happened. This class of error sits in more dashboards than I care to think about.
The last drawing is a counter image. First the truth, then the usual mistake.
The second mean is not computed wrongly. It is computed correctly over a sample the bad cases have quietly left. That is how you get a chart that reassures you while the stock burns down.
Six questions for any forecast
You do not have to build anything for this. Six questions do the job, and they work against any number somebody puts in front of you. Including mine.
- Does it name the split? How much of the spread is dice, how much is not knowing? Anyone who cannot answer has not computed it.
- Does it say what to measure next, and what that buys? A forecast without that sentence is a horoscope with decimal places.
- Does it carry its own error bar? And does the fixed number disappear when the bar gets too wide?
- Is there a dial for tail strength? Then throw it out. It exists so the result can look the way somebody wanted it to look.
- Does the comparison run on the same dice? Variant equals base must give exactly zero. Anyone not guaranteeing that is selling you noise as an effect.
- Is there a track record against the base rate? Replayed using only events dated strictly before the outcome. Otherwise the model is marking its own exam.
Question six is the hardest. A rare event scores beautifully under the prediction "never happens", which is why the reference is always the historical frequency. Score below it and you would have been better off quoting the base rate and going home. And if too many outcomes land in the tails of your own distributions, those distributions were too narrow. That is the classic failure, and exactly what this tool exists against.
When I advise against it
Not every decision needs this, and saying so is part of the thesis.
- When the decision comes out the same at any number. Then do not compute, decide. A simulation that cannot change anything is expensive reassurance.
- When you cannot decompose the system into drivers and relationships you can name. A model made of fog produces fog with decimal places.
- When you need the number to justify rather than to learn. A spreadsheet will do, and it is cheaper.
- When nobody is going to run the measurement the thing ends up recommending. Then the result is an observation about your organisation, not about the market.
What I do with it
I built the machine because I needed it for client cases and found nothing that delivers the split cleanly. It is called Balane Studio: a Python engine with a desktop interface, cases are local files, no cloud requirement. Modelling and deciding are free. Comparing and the track record sit behind a licence. Understanding costs nothing, proving does.
What I learned building it is independent of the tool, and it is the line I will close on:
Demand the split. Budget for measuring instead of more forecasting. And when somebody hands you a number without its own error bar, you did not get a forecast. You got an opinion dressed up as arithmetic.