The natural fluctuations of matter could help computers solve problems. But when a machine computes with chance, how do we know its answers are right?
Listen to this article
15:53
COMPUTING WITH CHANCE
Using nature’s randomness to help solve problems.
A p-bit is a small element that fluctuates between two states, with a probability we can influence. Connect many of them and their collective activity can be arranged to explore a probability distribution. Physical implementations can draw their randomness from fluctuations in electronic circuits or magnetic devices. [1]
This suggests a different approach to certain kinds of computation. A conventional computer can simulate the steps of a sampling algorithm. A purpose-built probabilistic machine aims to make the physical dynamics perform part of that work directly. The possibility of doing useful sampling with less energy is one reason this field is worth watching. [2]
Imagine being able to explore many possible configurations of a model at much lower cost. That could expand what is practical in probabilistic inference, scientific simulation and some forms of machine learning.
Turning that possibility into useful answers brings us to a question of trust: how do we know when the probabilities a machine produces faithfully describe the problem we gave it?
THE TRUST PROBLEM
Which probability are we estimating?
A one-in-a-hundred chance of losing a network connection and a one-in-a-million chance could lead to different decisions about paying for a backup. Calling both “unlikely” leaves out the difference that matters. A failure with a small probability of occurring is a rare event.
A computer can estimate the chance of connection loss by generating possible combinations of working and failed connections. Each combination is a sample. The fraction of samples that leaves a destination disconnected estimates the chance of connection loss.
For that estimate to be useful, the samples need to represent the model we gave the computer. Repeated runs can agree with each other even when the sampling process has changed the probabilities. Collecting more samples may make the answer more precise while leaving that mismatch in place.
We studied this problem using small software models whose intended answers we could calculate independently. We also calculated, where possible, the probabilities implied by the sampling rules themselves. Comparing those answers with the recorded results let us ask what the estimates were actually describing.
The experiment included two sets of checks. General checks examined overall patterns in the samples; event-specific checks examined observations of the particular outcome being estimated. A check gave “false assurance” when it approved an estimate that was wrong by more than the study’s allowed margin: below half or above twice the intended probability. The first example shows why even an event-specific check can miss that kind of error.
A PRECISE WRONG ANSWER
One in 100,000 became one in 739.
Our simplest model contained sixteen simulated arrows, each pointing up or down. Five particular arrows each had a one-in-ten chance of pointing up. If those choices are independent, all five point up together about once in 100,000 samples.
We deliberately gave simultaneous updates a shared source of randomness. Each arrow kept its individual chance of pointing up, but the arrows’ choices became related. Under the specified sampling rule, the calculated chance of all five pointing up together became about one in 739.
The 256 recorded runs all passed both sets of checks. They also produced estimates far from the intended one-in-100,000 answer. Yet their uncertainty intervals mostly described the changed process accurately: 243 of the 256 intervals included its calculated one-in-739 probability, while none included the intended probability.
An uncertainty interval describes a range around an estimate. Here, narrowing that range could not repair the difference between the model and the process generating the samples. The checks were reassuring about estimates of a probability the original model had never specified.
The same distinction appeared in the synthetic network example. Shared randomness preserved the availability of each connection individually, while making the calculated chance of at least one disconnected destination about five times higher. Knowing how each component behaves on its own can leave an important part of the system’s behavior unexplained.
The model behind this example
The five-arrow comparison is B1/S3/P4, with intended probability 0.00001 and ideal stationary probability approximately 0.00135257. The shared input has latent-Gaussian correlation 0.30; this is not the correlation between the arrows themselves. Each stored interval has a nominal 95% level. Their observed containment of a calculated reference does not establish general calibration.
The network is a generated graph with fourteen nodes, twenty connections and five destinations. Its event is at least one destination being unreachable from the source. The calculated event-probability increase is 4.96-fold. This is a connectivity model, with no packet traffic or measured outages.
WHEN ROUNDING MATTERS
Small settings can hold a model together.
Shared randomness changed how the arrows moved together. Rounding produced a different kind of change: it altered the model’s settings before sampling began.
In a model of 64 interacting arrows, the chance that at least 48 pointed up together was about one in 53. We rounded its settings to a grid spaced by 0.05. The interactions were small enough to round to zero, and the preference for one direction changed slightly. The resulting model described a much less likely event.
KNOWN PROBABILITY IN THE ORIGINAL MODEL1.8813%About 1 in 53
The chance that at least 48 of the 64 arrows point up.
MEDIAN ESTIMATE FROM THE ROUNDED MODEL0.0007%About 1 in 143,000
The middle of 256 separate simulation estimates.
The middle estimate across 256 runs made the event look roughly 2,700 times less likely than the original model specified. All 256 estimates fell outside the allowed range around the original answer, and all passed the general checks.
The event-specific checks withheld most approvals. The bar below accounts for every run: most lacked enough event evidence, nine failed another requirement, and nine passed despite being wrong against the original model.
238 · Not enough evidence to decideApproval withheld. The paper calls this “abstention.”
9 · Failed a checkA required check failed, so the estimate was not approved.
9 · Approved, but wrongThe checks passed despite an estimate outside the allowed error range.
Withholding an answer prevented many incorrect estimates from receiving approval. It did not restore the interactions lost through rounding, so the few approved answers still described the altered model. This is why translating a model into a device’s available settings is part of assessing accuracy. Extropic’s work on compiling stochastic programs likewise treats representation errors and the quantities a task needs as explicit concerns. [3]
How to read these numbers
The known probability is an independent calculation for the original model. The median is the middle of 256 recorded estimates after rounding; it is not another exact probability. This example uses the B2 positive event and sequential full sweeps, S1. Couplings of 1.35/64 rounded to zero, and bias changed from −0.04 to −0.05. The rounded model’s calculated probability is approximately 0.000006916. None of these runs had zero event hits. Reciprocal probabilities describe frequencies, not measured waiting times.
WHAT A LONGER RUN CHANGES
A run can pass through the right answer.
Once we knew that sampling rules could change the long-run probability, we asked what a finite run would report on its way there. The starting state can still influence the average estimate after many updates, so the result can depend on when we stop counting.
For one simultaneous-update setting in the 64-arrow model, the calculated average estimate began too high and eventually became too low. Between those extremes, it crossed the correct answer. At that point, an upward influence from initialization canceled a downward error in the long-run probability.
Points show calculated averages at selected runtimes, with connecting lines to guide the eye. The dashed line marks the intended probability. Crossing it means the average matches; it does not mean individual runs give dependable answers.
At the crossing, the predicted variation between runs was still about as large as the probability we wanted to estimate. A correct average could therefore hide a very uncertain individual answer. Accounting for both systematic error and variation showed that running longer initially helped, then hurt at a later evaluated runtime. A second event under the same rule kept improving over the evaluated range.
The useful question is therefore how more computation changes the accuracy of the particular answer we need. Simply collecting more samples does not tell us whether we are reducing uncertainty, approaching a biased answer, or temporarily balancing one error against another.
Calculation and uncertainty
This is B2/S3 with gain increased by 10%, 25,000 warmup sweeps, fair-spin initialization and eight ideal independent chains. The positive-event expectation matches its intended probability near 1,280,156 retained draws per chain; the predicted standard deviation is approximately 0.99 times that probability. These are finite-state calculations under the specified transition rule, not additional GPU runs or a validated stopping rule. The paper also reports an unresolved excess of observed variability in one partial-update condition and a separate underdispersion comparison.
WHAT APPROVAL SELECTS
The checks change which answers we see.
Runtime determines how much evidence a run collects. An approval rule then determines which of those runs we treat as usable. To isolate the effect of that choice, we returned to the rounded 64-arrow model, where the altered probability could be calculated and the full-update samples were independent.
Across two million samples, the event should appear about fourteen times on average. Consider a simple rule that accepts only completed runs with at least twenty observations of the event. Runs where chance produces more events are more likely to qualify. Under this rule, the calculated mean estimate among accepted runs is 55% above the rounded model’s own probability.
The same selection affects uncertainty intervals. Before selection, ordinary intervals labeled “95%” contain that probability in about 96% of hypothetical runs. Among runs that clear the twenty-hit threshold, only about 63% of the intervals contain it. Requiring a hit in each of eight separate sample sequences lowers that calculated share to about 58%.
The three copper points show calculated coverage under independent sampling from the rounded model. The separate purple square shows observed coverage among full-policy approvals: 15 of 33 saved intervals. Both comparisons use the rounded model’s probability; the count model does not reproduce the full approval procedure.
The purple square shows what happened under the full procedure used in the experiment, which had additional checks. Across the three full-update methods, it approved 33 of 768 runs. Fifteen of those intervals contained the rounded model’s probability, about 45%. The calculation and the recorded result are different comparisons, but both make approval part of the question: how reliable are the answers left after we decide which ones to accept?
What “coverage” means here
Coverage is the share of intervals containing the reference probability across repeated runs. The three copper points are finite-binomial calculations for ordinary Wilson intervals with a known sample count. The square summarizes saved intervals using estimated effective sample size under the full policy. The simple count rules do not reproduce that complete policy or establish a replacement interval method.
THE WIDER EXPERIMENT
Fewer wrong approvals, and fewer answers.
The complete experiment combined seven events, four sampling methods and five settings into 140 setups, each repeated 256 times. The general checks monitored energy, a summary of the model’s preferences and interactions, and magnetization, the overall balance between up and down arrows. The event-specific checks examined observations of the chosen outcome.
Each square below represents one setup. Copper means at least one wrong estimate received approval; matching positions refer to the same setup in both grids.
General checks
81 of 140 setups had at least one incorrect approval
Event-specific checks
50 of 140 setups had at least one incorrect approval
At least one wrong estimate approved
No wrong estimates approved
Copper can mean one wrong approval or many. Dark can include withheld decisions. These grids count affected setups, not the fraction of all estimates that were wrong.
Event-specific checks reduced the number of wrong approvals from 15,504 to 9,265. They also reduced total approvals from 35,344 to 25,750. Every event-specific approval in these data passed the general checks too, so part of the improvement came with returning fewer answers.
That tradeoff matters when choosing how to use a diagnostic. A rule can prevent some incorrect answers from being accepted while leaving the underlying model unchanged, or while changing the reliability of the answers it does accept. The paper reports approvals, withheld decisions and errors together so those outcomes can be distinguished.
How to read the grid
Rows show three independent-arrow events, two events in the 64-arrow model, agreement with a reference pattern in another interacting model, and synthetic network connection loss. The four five-column groups update arrows sequentially in random order, in conditionally independent groups, all together, or half together. Within each group, settings are unchanged, rounded, gain reduced, gain increased, and shared randomness introduced.
Comparing policies at the same approval count
Retaining conventional approvals uniformly at random within each setup until their counts match the event policy gives an expected 9,404.25 wrong approvals, compared with its observed 9,265. This descriptive comparator inherits how the event policy allocates approvals across setups; it does not uniquely identify the value of event information. Under a separate tenfold-error definition outside the 64-arrow benchmark, the event policy removes 3,487 approvals but only four errors, increasing the fraction of approved answers that are wrong. Both comparisons remain in the paper.
WHAT THESE EXAMPLES GIVE US
Reference answers for testing the whole calculation.
The value of these examples is that a wrong answer can be investigated. Comparing an estimate with the intended probability identifies an error. Calculating the probability generated by the update rule can explain its direction and scale. Finite-run calculations then show what initialization and uncertainty add, while selection calculations show what happens to the estimates that receive approval.
For 133 of the 140 setups, we could calculate a long-run probability for the specified sampling rule. In those setups, 9,009 approved estimates were below half or above twice the intended probability. Of those, 9,008 were within that broad margin of the sampling-rule probability instead. This is a broad comparison, but it helps explain why a reassuring diagnostic can accompany an answer to a changed probability problem.
These cases also suggest where additional checks could look. In the independent-arrow example, shared randomness preserved average energy and magnetization while changing their variation. In the interacting model, matching average energy with a single temperature-like adjustment still left event probabilities far apart. Checking how components behave together can therefore add information beyond checking each component or a single overall average.
Scope and unresolved questions
The study comprises 35,840 software runs on one NVIDIA GeForce RTX 3090 Ti. The later reference calculations and selected case studies are exploratory analyses of retained evidence. The underlying distinctions are established in approximate sampling and selective inference; this study supplies worked comparisons and reusable reference answers.
Seven stationary reference laws remain unavailable. The paper retains discrepancies between predicted and observed variability, including one partial-update excess and one full-update underdispersion comparison. It does not establish a new diagnostic or independently reconstruct every original trajectory.
Two original analysis procedures departed from the registered plan: handling zero-event estimates and assessing uncertainty across comparisons. A later conservative sensitivity analysis treated disputed classifications as unknown and retained positive differences favoring the event policy in 26 setups, under independent-repeat assumptions. Those results remain post hoc, alongside the original observations and their qualifications.
WHY IT MATTERS
Making cheaper samples useful.
The attraction of thermodynamic computing is the possibility of letting physical systems perform sampling that would otherwise consume substantial digital computation. A conventional processor could prepare a model, send part of the work to fluctuating circuits, and use the resulting samples in a larger calculation.
For that approach to help, the samples must support the answer the application needs. If a decision depends on several components taking particular states together, calibrating each component separately leaves their joint behavior to be checked. If a long run produces consistent estimates, those estimates still need a relevant probability reference. Earlier accelerator research already distinguishes convergence from accuracy; our cases make the consequences concrete for particular rare events. [4]
A natural hardware follow-up would start with small problems whose answers are known, calibrate individual responses, and then measure joint events under the device’s update schedules. The reference calculations could help distinguish a change in the represented model from finite sampling uncertainty or an effect of the reporting rule. Larger software experiments could examine how these distinctions change with scale, while physical measurements would test how well the assumed dynamics describe a real device.
The energy comparison belongs to that complete task. How much work does it take to reach a specified accuracy, including preparation, sampling, readout, checking and repeated attempts? This study does not measure a hardware energy advantage. It provides concrete probability-estimation cases with which to ask whether faster, cheaper samples lead to answers we can use.