Start with your documents
Ask one causal question about one bounded case. Upload several texts or embedded-text PDFs, then follow the evidence in the resulting report.
Process Tracing compares rival explanations against a source packet, shows which evidence distinguishes them, and leaves every conclusion traceable to the underlying text.
Open the ten-case study, inspect one focused comparison, enter either case, and open one exact source anchor. Finish with Record a research review. If time permits, the saved Portugal 1974 attempt shows how a run stops without manufacturing a conclusion.
Follow ten separately audited full case studies into five focused comparisons, then inspect the exact case surfaces and source anchors behind each finding. The study identifies bounded pathway variation; it does not estimate a general causal effect.
Need execution controls or stage-by-stage diagnostics? .
Continue an analysis or open its latest report. Earlier attempts are kept in run history.
Loading your analyses…
Ten purposively selected episodes were each studied before a separate five-block comparison. The result identifies differences in pathways and evidentiary limits; it is not a representative sample, effect estimate, or automatic historical verdict.
Loading the accepted study foundation…
A method for comparing causal explanations within a case while preserving the evidence, judgments, uncertainty, and claim limits behind the conclusion.
The workbench converts a bounded research question and source packet into an inspectable causal argument. It states rival explanations, evaluates the diagnostic value of source-grounded propositions, reconstructs the temporal mechanism, and subjects central judgments to separate audits. The architecture combines narrative Bayesian process tracing with explicit provenance, deterministic multi-hypothesis aggregation, dependence controls, and adversarial review. It also distinguishes the historical process that produced an outcome from the archival process that produced the evidence now available. The numerical Test stage is experimental: its LLM-generated values are algebraically coherent relative evidentiary scores, not empirically calibrated estimates of P(E | H). Its output is therefore normalized comparative evidentiary support—not a posterior probability, historical truth estimate, population effect, or substitute for source criticism.
Process tracing asks how an outcome came about inside a particular case. The object of inference is not a correlation between variables across many cases, but a causal explanation connected to observable events, decisions, constraints, and transmissions within one bounded historical record.
The project sits between two Bayesian traditions. Narrative or likelihoodist process tracing preserves historically specific explanations and asks how expected each observation would be under each rival. Formal model-based work declares a causal graph and derives implications from its structure and parameters. Both can use Bayes' rule, but they differ in what counts as a hypothesis, how mechanisms are represented, and where likelihood judgments come from. For a single source-rich case, this system is narrative-primary: it retains evidentiary texture while externalizing the reasoning that would otherwise remain implicit. Formal cross-case estimation is kept separate for settings in which cases share variables and the design can identify population or counterfactual quantities.
The system therefore does not ask a model for a free-form historical answer. It moves through five reviewable operations:
| Operation | Semantic work | Deterministic or independent control | Retained result |
|---|---|---|---|
| Bound | Interpret the question and source scope. | Validate the inference-design and source-packet contracts. | Frozen question, corpus, mode, and claim limits. |
| Specify | Construct mechanisms, rivals, and predictions. | Separately audit every rival relationship before confirmatory testing. | Versioned hypothesis partition and exposure lineage. |
| Test | Judge the rival-relative diagnostic value of contextualized propositions. | Derive coherent score ratios, approximate dependence, and separately audit every non-equal comparison. | Raw and effective experimental evidence-by-rival score matrices. |
| Trace | Reconstruct the sequence and proposed causal transmission. | Separately audit every stage and edge against accepted source spans. | Typed temporal mechanism with unresolved links. |
| Qualify | Synthesize the strongest warranted conclusion. | Calibrate verdict strength and verify central evidence anchors. | Report, sensitivity, nonclaims, and research agenda. |
Every run declares how explanations encounter evidence. In a theory-first design, rivals and observable predictions are fixed before evaluation evidence is examined. A discovery/evaluation split develops explanations from one source set and evaluates them against a disjoint held-out set. An exploratory design may use the full corpus to discover mechanisms and alternatives, but evidence used to formulate an explanation is not presented as independent numerical confirmation.
The confirmatory gate examines every pair of rivals for overlap, complementarity, or absorption. Numerical comparative support is restricted to a defensible local mutually exclusive partition. Compatible, conjunctural, sequential, nested, and equifinal components remain visible but are not forced into winner-take-all percentages.
The target numerical estimand is deliberately narrow: posterior odds over a finite set of mutually exclusive, mechanism-specific explanations, conditional on the admitted corpus, declared priors, and calibrated diagnostic likelihoods. The current implementation has not earned that probability interpretation. It reports normalized comparative evidentiary support derived from experimental model scores. These scores rank how diagnostically favorable an item appears under each rival; algebraic coherence does not establish that they are proportional to P(E | H).
The current pipeline also includes a flat residual label for “none of the specified stories.” That is an open-world warning, not yet a genuine Bayesian hypothesis: without a predictive distribution for the heterogeneous alternatives it contains, it cannot compete on the same numerical footing. A future version will either model an imprecise mixture over named residual subtypes or exclude the residual from normalization and state that support is strictly conditional on the declared rivals.
Exposure changes the claim. If an observation helped generate a hypothesis, treating the same observation as independent confirmation would count it twice. The system currently records evidence visible during hypothesis generation and sources informing the prior. Theory-first and genuinely held-out evidence can support confirmatory updating; full-corpus exploratory evidence remains available for explanation-building and mechanism appraisal but is not relabeled as independent support. This is a conservative workbench design policy, not the uniquely correct interpretation of iterative Bayesian inquiry. A future version will extend lineage to acquisition, retrieval queries, inclusion and exclusion, extraction, outcome definition, observable-prediction construction, and prior formation.
The present rival audit detects overlap and incompatible relationships but does not yet make support invariant to how explanations are split or merged. Dividing one parent explanation into two mechanism variants can redistribute normalized mass even when the substantive evidence has not changed. Plan 043 therefore adds a hierarchical parent → mechanism-variant representation, parent-level reporting, and mandatory split/merge sensitivity.
A likelihood judgment about evidence rests on more than a story about what caused the historical outcome. It also depends on how the alleged event became an observable trace. The method therefore distinguishes two models that are often collapsed:
This distinction matters because a trace can be present for at least two reasons: the underlying event occurred and was recorded, or the claim was produced even though the event did not occur as stated. Propaganda, bureaucratic formulae, retrospective rationalization, genre conventions, selective survival, and extraction error create false-positive channels. Conversely, censorship, secrecy, destruction, and archival absence can suppress traces of events that did occur.
Retaining H in both trace channels is essential: even when the underlying fact is held fixed, different explanations can imply different incentives to fabricate, suppress, classify, standardize, or preserve its record. The current workbench captures source provenance, coverage, dependence lineage, and source-silence opportunity, but it does not yet represent this complete typed trace-production model. That is an explicit future-version target, not a capability implied by the present interface.
Every evidentiary proposition must resolve to accepted source text and stable provenance. The exact quotation is a citation anchor, not the whole evidentiary object. Meaning can depend on the surrounding paragraph, genre, authorship, audience, purpose, competence, authenticity, temporal proximity, translation, editorial history, and incentives to misrepresent. The current exact-anchor requirement provides strong provenance but only partial contextual source criticism; Plan 043 makes contextual windows and those source fields mandatory for update-bearing evidence.
For each evidence item, the testing stage proposes one positive vector of relative evidentiary scores across all rivals. Pairwise ratios are derived from that common vector rather than elicited separately. This matters when there are three or more hypotheses: independent pairwise judgments can violate reciprocity or transitivity. Equal vector components mean the model judges the evidence nondiscriminating between those rivals.
This construction establishes algebraic coherence. It does not establish calibration: q may be a stable model score without being proportional to P(E | H). Paci's 2026 benchmark found that GPT judgments could be internally coherent and run-stable while remaining substantially different from expert estimates, predominantly overweighting evidence and sometimes changing the leading hypothesis. Automated magnitude estimation is therefore the principal validity bottleneck.
Priors remain conceptually distinct from evidence scores. Typed prior specifications retain their method, weights, rationale, and source provenance so the same source is not silently used twice. The current deterministic layer aggregates centered scores and exposes sensitivity:
The resulting weights are normalized comparative evidentiary support within the declared rival set. They are not posterior probabilities or posterior odds. Legacy internal fields still use likelihood/posterior names; those implementation names do not license the stronger public interpretation.
Hoop, smoking-gun, doubly-decisive, and straw-in-the-wind tests summarize the shape of a likelihood comparison. They do not replace the likelihood model. A hoop test describes evidence strongly expected if a hypothesis is correct; failure is damaging, while passage may be weak support. A smoking-gun trace is unlikely under rivals; its presence is strongly supportive, while its absence may say little. The labels make diagnostic logic readable, while the multi-rival vector carries the coherent comparison.
The system retains the originally elicited vector and records each deterministic transformation separately: relevance adjustment, dependence pooling, eligibility decisions, caps, and sensitivity perturbations. No workload cutoff or display convention is treated as a law of process tracing. A conclusion that changes under plausible priors, likelihood magnitudes, or guardrail settings is reported as fragile.
A valid joint likelihood generally factors conditionally:
The current system instead pools evidence through a claim → event → source → author hierarchy. This reduces obvious duplication, but a scalar discount is not equivalent to a joint likelihood model. Dependence may also be hypothesis-specific—for example, two documents may share one propaganda origin under one rival but represent separate observation channels under another. Plan 043 replaces the hierarchy as the sole representation with an evidentiary-ancestry graph, joint cluster judgments where needed, hypothesis-specific dependence, and leave-one-cluster-out sensitivity.
Source silence currently enters numerical aggregation only when an admitted source was positioned to record the predicted trace, the search is complete, no partial observation exists, and a separate audit accepts that opportunity-to-observe judgment. Missing source classes limit the conclusion; they do not prove absence in the world. A fuller trace-production model will replace the binary opportunity gate with rival-relative probabilities of nonappearance while preserving qualitative fallback when observability cannot be defended.
Exact quotations do not prevent selection on diagnosticity or latent model knowledge. A future version will retain the complete candidate-evidence universe—including rejected and nondiagnostic passages—freeze the retrieval and stopping rule, and audit inclusion independently of score extremity. It will also compare famous-name packets with anonymized entities and dates, fictional substitutions, permuted rival labels, and counterfactual packet variants. Material score changes caused by “Napoleon” rather than the admitted evidence violate the intended corpus-conditional claim.
The method constructs a forward temporal graph from source-grounded stages. Each transition is classified as observed sequence, supported causation, contested causation, or unresolved. Temporal order alone never certifies causation. An unresolved link must identify the evidence that would clarify the causal break.
This distinction prevents a fluent narrative from hiding inferential gaps. A source may establish that event A preceded event B while leaving open whether A caused B, whether a third process produced both, or whether the relevant transmission is missing from the admitted record.
The graph is a qualitative causal audit, not a machine for reading likelihood magnitudes from topology. Without defensible parameters, a graph can reveal a missing pathway, confound, temporal contradiction, or translation loss; it cannot determine whether a likelihood ratio should be two or five. The narrative estimator supplies the magnitude, while the graph forces the causal and evidentiary structure behind that judgment into inspectable form.
The current category supported causation is still a corpus-conditional model judgment. Plan 043 will require each promoted edge to face a miniature rival-relative test: direct transmission versus a plausible alternative transmission, a common cause, and no causal connection. Until that control exists, the graph should be read as a typed and separately reviewed causal argument—not as independently identified causal structure.
The component that produces a judgment does not certify it. Separate review stages inspect the rival partition, every materially discriminating likelihood comparison, the mechanism stages and edges, and the central claims in the final synthesis. Reviews are bound to the exact candidate artifacts and accepted source spans.
Rejected judgments, corrections, attempts, artifact hashes, model routes, costs, and final outputs remain inspectable. This creates traceability of one analytical execution; it does not by itself establish rerun reproducibility, because proprietary model versions, stochastic inference, retrieval indices, and tools can change. Future artifacts will bind exact model and version, prompts, reasoning and decoding settings, tool/runtime versions, source hashes, retrieval-index version, and seed where supported, then measure repeated-run variation.
Separate prompts and blinded roles provide procedural separation by reducing anchoring and self-certification. Different model families can add model independence. Epistemic independence—different information or differently trained priors—is much harder: several models may reproduce the same historiographic stereotype. The default label is therefore separate audit. “Independent audit” is reserved for explicitly blinded cross-family or human review with recorded independence provenance. Agreement is supporting evidence, not proof, and same-model repetition is never described as independent corroboration.
Uncertainty enters through the source packet, rival partition, priors, diagnostic likelihoods, dependence assumptions, mechanism translation, and missingness judgments. The report keeps these sources separate rather than compressing them into one generic confidence score. It emphasizes whole-number comparative support, sensitivity ranges, rank stability, contested mechanism edges, and explicit source gaps.
A high comparative weight can coexist with weak evidentiary resolution when every rival is poorly observed or when one explanation merely loses less badly. Conversely, an indeterminate conclusion can be methodologically successful when the admitted record does not discriminate. Contract completeness, historical argument quality, and evidentiary resolution are separate evaluations.
Equations, graphs, and audit logs can lend opaque judgments a false aura of objectivity. The architecture is answerable only if it can be tested against a simpler narrative-only baseline. Representative validation should measure calibration, discrimination, dependence detection, duplicate-overcounting control, rank accuracy, and expert contestability on held-out or adversarial cases. Thresholds and scoring rules must be fixed before evaluation; otherwise procedural sophistication becomes decoration.
The validation design has two strata. Synthetic and semi-synthetic cases expose a known data-generating process, allowing objective tests of calibration, rank recovery, duplication, dependence, trace production, and missingness. Real historical cases use blinded expert adjudication to evaluate defensibility, context use, error discovery, agreement, and useful disagreement—not an imaginary gold-standard historical posterior.
The benchmark must be able to reject numerical use. The post-MVP Plan 043 design predeclares tolerances for magnitude bias, rank error, proper-scoring performance, anonymization and label-permutation sensitivity, split/merge instability, dependent-cluster influence, contextual reversals, and rerun variation. Its corrected full benchmark remains unexecuted and is not an MVP release criterion. If future calibrated automation does not beat a simpler qualitative baseline on held-out evidence—or if these stress tests materially change the leading conclusion—the default product will suppress normalized numbers and retain scoring only as a structured deliberation aid.
On 17 August 2026, the hosted workbench received four previously uninstalled documents about Portugal's 25 April 1974 regime collapse. Eight accepted Gemini provider calls completed extraction through the experimental update for $0.110634; a ninth call stopped at mechanism reconstruction when the provider returned HTTP 503. The attempt was retained without a model switch or substantive retry, so it produced no final causal narrative.
The partial result is methodologically informative. Nine propositions from all four sources resolved to exact admitted quotations. Because the full corpus also generated the four candidate explanations, every proposition was excluded from diagnostic updating and all five displayed weights—including the residual—remained at their equal starting value. That equality is not a finding. The rival audit also classified all six pairs as potentially complementary and two pairs as absorptive, even while allowing the exploratory workflow to continue because each pair had an opposed prediction. This is direct implementation evidence that the present exploratory path can generate a plausible research agenda but does not yet implement the target typed relationships among compatible, conjunctural, nested, and sequential mechanisms.
A completed run establishes what the admitted corpus supports, weakens, or leaves unresolved; how the declared rivals compare under the accepted diagnostic judgments; which causal links remain contested; and what evidence would most improve the inquiry.
Process tracing remains a within-case method. Cross-case inference therefore begins only after each case has a separately inspectable full causal study and then applies a separate selection and comparison design. A compact trace is a useful intermediate projection, but it does not replace the rival specification, evidence testing, mechanism qualification, conclusion, and uncertainty analysis required of a completed case. A common discovery corpus may orient case selection, but it is not reused as independent case evidence.
The current study asks: across bounded revolutionary and comparison episodes, what combinations and sequences distinguish episodes that establish a materially different governing order from episodes that do not? The unit is one connected challenge to an identified incumbent sovereign order with an observable mobilization, response, and endpoint. “Transformation” means that challengers displaced that order and established materially different governing authority within the frozen episode boundary. Mobilization, violence, leadership turnover, territorial control, or the label revolution is not sufficient.
The eligible universe was first assembled from three declared English-language discovery surfaces and reconciled into 1,963 provisional episodes. Twenty finalists then faced eligibility and pre-outcome comparability review before the separately retained endpoint ledger was joined. The final five–five balance is purposive contrast, not genuine outcome blinding, random sampling, prevalence, or a population estimate: names and boundary descriptions could reveal outcomes, and source availability was a feasibility constraint rather than a reason to select a famous case.
Selection follows structured, focused and theory-relative logic. Cases fill declared roles—anchor pathway, similar-context opposite outcome, same-polity temporal contrast, alternative path, external intervention, near miss, deviant non-transformation, and construct boundary—rather than pretending that one generic “most similar” rule solves every inferential problem. Replacements must fill the same role and preserve the frozen outcome and block design; inconvenient cases cannot be swapped merely for richer or easier sources.
| Focused block | Cases | Inferential use and limit |
|---|---|---|
| Anchor opposite pathway | Romania 1989; Myanmar 1988 | Stable opposite-endpoint contrast in mobilization, coercion, elite alignment, and authority transfer; not an effect estimate. |
| Russian sequence | Russia 1905; February 1917 | Same-polity contrast in concession, coercive cohesion, organization, and displacement; shared institutions and historical sequence make the cases dependent. |
| Late Cold War contrasts | Romania; Myanmar; Hungary 1956; East Germany 1953; Velvet Revolution | Compares domestic alignment, external coercion, information, restoration, and displacement while retaining bloc structure, diffusion, and historiographic dependence. |
| Transformation and boundary scope | French Revolution; Iranian Revolution; Velvet Revolution; May 1968 | Probes different transformation configurations. May 1968 is used only to test the construct boundary, not as substantive evidence for a general protest mechanism. |
| Non-transformation deviance | Myanmar; Russia 1905; Hungary 1956; East Germany 1953 | Asks whether serious but restored challenges strain the same explanation in different ways; no case scores are averaged. |
Within each completed case, rivals are tested against source evidence and the temporal mechanism is qualified. Across cases, the system compares only declared constructs, configurations, and sequences carried by those full studies. It does not pool case-level support weights, create one ten-case rival simplex, count overlapping blocks as independent observations, or infer necessity and sufficiency from recurrence. Every cross-case observation must step back to named cases, then to their exact evidence anchors; dependence and construct failures can exclude a case from a particular claim without removing its value as a boundary test.
The current implementation has ten completed source-limited full case studies and a separate full 18 Brumaire depth reference. Its five-block synthesis was rerun from those ten terminal case artifacts and separately audited. The result remains source-limited: it is a bounded comparative interpretation of the installed corpus, not a population estimate or a general theory of revolution.
When many cases are genuinely comparable, the project has a separate formal path built around explicit causal models and CausalQueries. That path can target identified counterfactual or population estimands only when common variables, case comparability, measurement, and model assumptions justify them. A single case is never sent through both engines and presented as if narrative support and a population causal effect were the same quantity.
The public method distinguishes demonstrated controls from the quality-optimal architecture that guides future work. “Implemented” means the software path exists and has focused technical evidence; it does not mean that substantive historical validity has been established.
| Capability | Status | Current boundary |
|---|---|---|
| Inference-design and exposure lineage | Implemented | Public output changes with theory-first, held-out, or exploratory status; exploratory discovery is not presented as independent confirmation. |
| Coherent multi-rival scoring | Implemented | One positive score vector per item, deterministic log-space aggregation, typed priors, and sensitivity readout. Calibration as P(E | H) is not established. |
| Separate discriminator and mechanism audits | Implemented | Complete typed review with exact-quote and artifact-hash binding; cross-family or epistemic independence is not guaranteed. |
| Dependence and source silence | Partial | Hierarchical pooling and strictly gated source-silence logic exist; pooling is an approximation and dependence remains scalar rather than hypothesis-specific. |
| MVP prompt-behavior smoke | Diagnostic only | Eight obvious behaviors appeared plausible, but the label-permutation retry changed a unique lead into an audited tie. Numerical support is therefore withheld from the MVP headline; retained scores remain inspectable and do not establish calibration. |
| Score uncertainty and calibration | Plan 043 | Point judgments, prior/driver sensitivity, and rank stability are exposed; held-out calibration, Paci-compatible comparison, bands, and joint propagation remain unfinished. |
| Context, selection, and contamination controls | Plan 043 | Exact anchors and source provenance exist; contextual packets, complete candidate custody, stopping rules, and anonymization/counterfactual probes are planned. |
| Partition, residual, and mechanism-edge sensitivity | Plan 043 | Rival relationships and edge categories exist; hierarchical split/merge tests, a predictive residual, and local edge alternatives remain planned. |
| Trace-production model | Plan 043 | Provenance and coverage are retained, but causal production, recording, survival, hypothesis-specific incentives, and false-positive channels are not yet a complete typed model. |
| Representative validity benchmark | Plan 043 | Technical integrity and demonstrations exist; known-DGP calibration, blinded historical review, stress tests, and decision sign-off remain to be earned. |
The 18 Brumaire report compares three predeclared explanations across 82 evidence records and retains its fragile lead, contested links, exact evidence, and next-source agenda.
Define one causal question and supply documents for one bounded case. The analysis preserves the corpus, considers rival explanations, and keeps every conclusion connected to its source.
These four decisions define the research design.
| Stage | Module | Input | Output | LLM? |
|---|---|---|---|---|
| 1 · Extract | pass_extract.py | Source text | Evidence items, actors, events, mechanisms, causal edges | Yes |
| 2 · Hypothesize | pass_hypothesize.py | Extraction | Competing causal hypotheses with observable predictions | Yes |
| 2.5 · Partition | pass_partition.py | Hypotheses | Accepted rival set and opposed prediction contrasts | Yes |
| 3 · Test | pass_test.py | Extraction + hypotheses | Experimental relative-score vector per evidence item; rejected discriminators trigger bounded re-elicitation | Yes |
| 3a · Audit splits | pass_discriminator_audit.py | Exact quotes + proposed score differences + rival predictions | Accepted/rejected judgments and bounded re-elicitation history | Yes |
| 3b · Absence | pass_absence.py | Hypotheses + extraction | Missing predicted evidence, source-opportunity judgment, and qualitative fallback when numerical admission fails | Yes |
| 4 · Aggregate support | bayesian.py | Effective score matrix + priors | Normalized comparative evidentiary support, sensitivity bands, and robustness labels | No — pure math |
| 4.5 · Mechanism | pass_mechanism.py | Extraction + rivals + accepted testing | Temporally ordered stages and status-bearing causal transitions | Yes |
| 5 · Synthesize | pass_synthesize.py | All prior outputs | Written narrative, per-hypothesis verdicts, steelman cases | Yes |
| 6 · Refine | pass_refine.py | Source text + first-pass results | Delta (new evidence, reinterpretations) → re-runs passes 3–5 | Yes (optional) |
Follow the analysis from source extraction through reviewed publication. Open a stage for its status and result.
What this does: Reads the bounded source corpus through typed extraction passes and retains historically observable items — evidence quotes, named actors and their roles, discrete events with dates, causal mechanisms, and directed causal edges. Each evidence item gets a relevance score and a diagnostic type (hoop / smoking gun / doubly decisive / straw in the wind).
Source fidelity rule: every evidence item must quote or closely paraphrase the input text. No hallucinated evidence. Circular evidence (derived from interpretation) is flagged.
| ID | Source text | Type | Relevance |
|---|---|---|---|
| evi_murat_grenadiers | 19 Brumaire: Murat entered the Orangerie and announced that the Council was dissolved | Smoking gun | 0.92 |
| evi_sieyes_director | May 1799: Sieyès elected to Directory after years of republican constitutional drafting | Straw | 0.88 |
| evi_first_consul_powers | Dec 1799: First Consul promulgated laws; other consuls advisory only | Hoop | 0.95 |
| evi_conspirators_rue_victoire | Evening 17 Brumaire: Conspirators assembled at rue de la Victoire | Smoking gun | 0.94 |
| evi_ancient_regime_1789 | 1789: Estates-General convened; Third Estate declared National Assembly | Straw | 0.22 (below threshold) |
What this does: Generates 3–6 competing causal hypotheses from the extraction. Each must specify a causal mechanism (not just an outcome description), name at least one observable prediction, and be genuinely rivalrous — if two hypotheses predict the same evidence in the same direction, they must merge.
Guardrails: anti-tautology, anti-circularity, a shared pinned outcome, distinct mechanisms, observable predictions, and system-recorded formulation exposure.
| ID | Label | Mechanism |
|---|---|---|
| h1 | Civilian revision | Sieyès-led constitutional redesign; Bonaparte was the military guarantor Sieyès needed |
| h2 | Military conversion | Bonaparte's personalist military authority converted directly into political dominance via 18 Brumaire |
| h3 | Path closure | Contingent institutional failures progressively closed alternatives until coup was only viable exit |
| h4 | Legal engineering | Elite legal-engineering by experienced jurists produced constitutional institutionalization of personal rule |
| h5 | Popular endorsement | Broad popular demand for order and stability provided the legitimating base for Bonapartist rule |
Decision: Do the hypotheses form genuine rival causal configurations with concrete opposed predictions?
Accepted fixture partition with typed opposed predictions.
What this does: For each evidence item, asks the LLM: "How likely is this evidence under each hypothesis, relative to the others?" Returns a likelihood vector, not independent pairwise scores — so ratios between hypotheses are coherent by construction.
Relevance gate: items with relevance < 0.4 are forced to LR = 1.0 (uninformative). Relevance = min(temporal, causal-domain). Cluster detection groups dependent evidence (e.g. all evidence from a single chain of events).
Loading…
Decision: Does the exact quote make the favored account materially more likely than its named rival on the accepted opposed prediction?
| Evidence / quote | Rival contrast | Proposed split | Judgment | Reason |
|---|---|---|---|---|
| No live audit loaded. | ||||
What this does: Evaluates what predicted evidence is missing from the text. If h5 predicts popular demonstrations of support but the text contains none, that is a failed hoop test. Each absence finding is rated "damaging", "notable", or "minor" with a reasoning note about whether the text would contain this evidence if it existed.
Design constraint: absence findings are qualitative only. They inform the synthesis narrative but do not receive speculative LR values — which would add noise to the Bayesian update.
| Hypothesis | Missing prediction | Severity | Reasoning |
|---|---|---|---|
| h5 · Popular endorsement | Popular demonstrations of active support for Bonaparte | Damaging | A text this detailed about the coup would mention mass rallies if they occurred; their absence is informative. |
| h1 · Civilian revision | Sieyès drafting sessions with jurists post-coup | Notable | Post-coup constitution was drafted quickly; evidence of Sieyès-led sessions may exist in sources not covered. |
| h3 · Path closure | Documented attempts at non-coup constitutional reform in 1799 | Notable | Text focuses on coup itself; reform attempts might appear in broader political history sources. |
What this does: Applies a coherent joint Bayesian update in log space: post_i = softmax(log prior_i + Σ log LR_i). Order-invariant. Per-evidence LRs are derived from the likelihood vectors as relative_likelihood / geomean(vector), so pairwise ratios are coherent by construction.
Then runs sensitivity analysis: for each hypothesis, perturbs its top 3 driver LRs ±50% and reports the posterior range. Robustness classification (robust / moderate / fragile) is mechanical — not LLM judgment.
Shaded band = posterior range under ±50% perturbation of top LR drivers. Not a probability of truth — a comparative ranking.
What this does: Converts the accepted evidence into a forward-only temporal DAG, then separately grades every transition as observed sequence, supported causal link, contested causal link, or unresolved link.
Reading rule: left-to-right order is necessary for causation, not sufficient. Evidence IDs ground what is observed; unresolved arrows name the next trace needed.
The live result renders here with a canonical edge table. Isolated nodes are invalid at this stage.
What this does: A separate semantic role checks every stage and transition against the accepted corpus. It may accept or weaken an edge, but it cannot strengthen one.
Failure rule: a material omitted stage or transition triggers bounded full-DAG replacement and re-audit, then blocks synthesis if unresolved.
The live resolution renders here as audit lineage plus stage, edge, correction, and omission tables. The raw typed contract remains available for inspection.
What this does: Writes the analytical narrative. The LLM has access to all evidence, hypotheses, LR vectors, posteriors, sensitivity ranges, and absence findings. It produces: a synthesis paragraph, a verdict for each hypothesis ("supported / partially supported / not supported / eliminated"), and a mandatory steelman case for every hypothesis — even eliminated ones.
Verdict calibration: verdict_calibration.py deterministically downgrades any verdict label that overstates the computed comparative support. The LLM writes reasoning; it does not have final authority on verdict labels.
The evidence overwhelmingly supports h2 (Bonaparte's personalist military conversion) as the decisive causal mechanism of 18 Brumaire. The concentration of military command, the rue de la Victoire conspiracy, and Murat's use of grenadiers to dissolve the Council of Five Hundred all point to deliberate military agency rather than civilian constitutional revision. However, this result is classified fragile: it rests on a broad accumulation of weak-to-moderate LRs rather than a few decisive smoking-gun tests. The hypothesis cannot be considered settled absent higher-diagnosticity primary sources documenting Bonaparte's intent and planning.
h4 (Legal engineering) receives marginal support: the constitutional architecture of Year VIII does show elite legal craftsmanship, but this appears to be a consequence of the coup rather than its cause. h1 (Civilian revision) has a non-trivial sensitivity range (up to 0.44 under perturbation) and should not be dismissed without stronger discriminatory tests.
| Hypothesis | Verdict | Posterior | Robustness |
|---|---|---|---|
| h2 · Military conversion | Supported (fragile) | 0.9942 | Fragile |
| h1 · Civilian revision | Partially supported | 0.0047 | Fragile |
| h4 · Legal engineering | Not supported | 0.0008 | Moderate |
| h3 · Path closure | Not supported | 0.0003 | Fragile |
| h5 · Popular endorsement | Eliminated | <0.0001 | Robust |
What this does: Second reading of the source text with the full first-pass context loaded. The LLM produces a structured delta: new evidence items it missed on the first pass, reinterpretations of existing items, spurious items to remove, and hypothesis refinements. Then passes 3–5 re-run on the updated extraction.
Pre-refinement artifacts are saved to pre_refine/ before being overwritten, enabling the Delta Board in the workbench to show before/after posterior comparison.
What this does: inventories terminal claims, checks each atomic claim for source or artifact entailment, and publishes the final result only when the complete review is accepted.
When the persisted repair limit is one, a blocked review becomes exact constraints for one replacement mechanism, independent re-audit, and revised synthesis. A second block stops permanently; extraction, hypotheses, testing, and numerical support never change.
What this does: Serialises the full ProcessTracingResult to result.json and renders a self-contained report.html. The report leads with the typed temporal mechanism DAG, then shows the evidence × hypothesis matrix, comparative support with sensitivity ranges, steelman cases, absence findings, provenance, source coverage, and hierarchical dependence.
No LLM calls — fully deterministic from the upstream structured results. Run again at any time to regenerate without re-running any LLM passes.
| Field | Value |
|---|---|
| run_id | run_20260624_200616_d4f6 |
| evidence_items | 59 (6 below relevance threshold) |
| hypotheses | 5 |
| dominant_hypothesis | h2 · Military conversion · posterior 0.9942 (fragile) |
| absence_findings | 3 (1 damaging, 2 notable) |
| dependence_clusters | 1 (Cluster A · 8 items · strength 0.70) |
| source_markers | 6 (A–F) · 49/59 evidence items traceable |
| historical_fixture_model | gemini/gemini-2.5-flash |
| historical_fixture_cost_usd | ~$0.08 |
A real grounded-theory incident and a real process-tracing evidence item consumed through the same source-identity contract.
Loading the hash-bound comparison…
See an authentic grounded-theory category revision beside process-tracing evidence that challenges a causal hypothesis.
Loading the hash-bound comparison…
See what current evidence suggests, what remains uncertain, and what would test an explanation more convincingly.
Loading the P5 pre-run acquisition agenda…
Each edge labelled with the Pydantic type that crosses it. LLM / Math / IO badges per node.
All data that flows between stages is typed. Every seam is a Pydantic model with Field(description=...) on every field.
Call order, LLM vs pure-math vs IO, and what crosses each boundary. Refine sub-sequence shown at bottom.