Process Tracing

See how a causal explanation holds up against the evidence.

Process Tracing compares rival explanations against a source packet, shows which evidence distinguishes them, and leaves every conclusion traceable to the underlying text.

Other ways to explore

Start with your documents

Ask one causal question about one bounded case. Upload several texts or embedded-text PDFs, then follow the evidence in the resulting report.

Read the completed case study

Inspect the theory-first 18 Brumaire analysis from rival explanations to exact source passages.

Understand the method

See what the system tests, what its evidence can establish, and where its claims stop.

Need execution controls or stage-by-stage diagnostics? .

Process Tracing Workbench · New analysis
Completed ten-case comparative study · source-limited

What distinguishes revolutionary transformation?

Ten purposively selected episodes were each studied before a separate five-block comparison. The result identifies differences in pathways and evidentiary limits; it is not a representative sample, effect estimate, or automatic historical verdict.

Loading the accepted study artifacts…

Loading the accepted study foundation…

Methodological note

Automated, source-auditable process tracing

A method for comparing causal explanations within a case while preserving the evidence, judgments, uncertainty, and claim limits behind the conclusion.

Within-case causal inferenceAI-assistedSource-steppable

Abstract

The workbench converts a bounded research question and source packet into an inspectable causal argument. It states rival explanations, evaluates the diagnostic value of source-grounded propositions, reconstructs the temporal mechanism, and subjects central judgments to separate audits. The architecture combines narrative Bayesian process tracing with explicit provenance, deterministic multi-hypothesis aggregation, dependence controls, and adversarial review. It also distinguishes the historical process that produced an outcome from the archival process that produced the evidence now available. The numerical Test stage is experimental: its LLM-generated values are algebraically coherent relative evidentiary scores, not empirically calibrated estimates of P(E | H). Its output is therefore normalized comparative evidentiary support—not a posterior probability, historical truth estimate, population effect, or substitute for source criticism.

Numerical validity status: experimental. The source trace, audit trail, and deterministic aggregation are implemented. Calibration of LLM score magnitudes as Bayesian likelihoods is not. Plan 043 is the binding gate for any stronger numerical claim.

1. The inferential task

Process tracing asks how an outcome came about inside a particular case. The object of inference is not a correlation between variables across many cases, but a causal explanation connected to observable events, decisions, constraints, and transmissions within one bounded historical record.

The project sits between two Bayesian traditions. Narrative or likelihoodist process tracing preserves historically specific explanations and asks how expected each observation would be under each rival. Formal model-based work declares a causal graph and derives implications from its structure and parameters. Both can use Bayes' rule, but they differ in what counts as a hypothesis, how mechanisms are represented, and where likelihood judgments come from. For a single source-rich case, this system is narrative-primary: it retains evidentiary texture while externalizing the reasoning that would otherwise remain implicit. Formal cross-case estimation is kept separate for settings in which cases share variables and the design can identify population or counterfactual quantities.

Core thesis. Language models do not make historical inference reliable by themselves. They make it feasible to decompose expert process-tracing labor into explicit, contestable operations—and to repeat audits, sensitivity checks, and provenance inspection at a scale that would otherwise be prohibitive.

The system therefore does not ask a model for a free-form historical answer. It moves through five reviewable operations:

BoundQuestion, case, outcome, and sources
SpecifyRivals, mechanisms, and predictions
TestExact evidence against each rival
TraceSequence and causal transitions
QualifyFinding, limits, and next evidence
OperationSemantic workDeterministic or independent controlRetained result
BoundInterpret the question and source scope.Validate the inference-design and source-packet contracts.Frozen question, corpus, mode, and claim limits.
SpecifyConstruct mechanisms, rivals, and predictions.Separately audit every rival relationship before confirmatory testing.Versioned hypothesis partition and exposure lineage.
TestJudge the rival-relative diagnostic value of contextualized propositions.Derive coherent score ratios, approximate dependence, and separately audit every non-equal comparison.Raw and effective experimental evidence-by-rival score matrices.
TraceReconstruct the sequence and proposed causal transmission.Separately audit every stage and edge against accepted source spans.Typed temporal mechanism with unresolved links.
QualifySynthesize the strongest warranted conclusion.Calibrate verdict strength and verify central evidence anchors.Report, sensitivity, nonclaims, and research agenda.

2. Research design, estimand, and rival explanations

Every run declares how explanations encounter evidence. In a theory-first design, rivals and observable predictions are fixed before evaluation evidence is examined. A discovery/evaluation split develops explanations from one source set and evaluates them against a disjoint held-out set. An exploratory design may use the full corpus to discover mechanisms and alternatives, but evidence used to formulate an explanation is not presented as independent numerical confirmation.

Rival explanation
A distinct account of the outcome that specifies a causal mechanism and at least one observation expected to differ from another account.
Observable prediction
A trace that should be more or less likely to appear if an explanation is correct, stated before that trace is used diagnostically.
Admitted evidence
Source content within the declared corpus whose text, identity, provenance, and role in the analysis are retained.

The confirmatory gate examines every pair of rivals for overlap, complementarity, or absorption. Numerical comparative support is restricted to a defensible local mutually exclusive partition. Compatible, conjunctural, sequential, nested, and equifinal components remain visible but are not forced into winner-take-all percentages.

The target numerical estimand is deliberately narrow: posterior odds over a finite set of mutually exclusive, mechanism-specific explanations, conditional on the admitted corpus, declared priors, and calibrated diagnostic likelihoods. The current implementation has not earned that probability interpretation. It reports normalized comparative evidentiary support derived from experimental model scores. These scores rank how diagnostically favorable an item appears under each rival; algebraic coherence does not establish that they are proportional to P(E | H).

The current pipeline also includes a flat residual label for “none of the specified stories.” That is an open-world warning, not yet a genuine Bayesian hypothesis: without a predictive distribution for the heterogeneous alternatives it contains, it cannot compete on the same numerical footing. A future version will either model an imprecise mixture over named residual subtypes or exclude the residual from normalization and state that support is strictly conditional on the declared rivals.

Selection and prior provenance

Exposure changes the claim. If an observation helped generate a hypothesis, treating the same observation as independent confirmation would count it twice. The system currently records evidence visible during hypothesis generation and sources informing the prior. Theory-first and genuinely held-out evidence can support confirmatory updating; full-corpus exploratory evidence remains available for explanation-building and mechanism appraisal but is not relabeled as independent support. This is a conservative workbench design policy, not the uniquely correct interpretation of iterative Bayesian inquiry. A future version will extend lineage to acquisition, retrieval queries, inclusion and exclusion, extraction, outcome definition, observable-prediction construction, and prior formation.

Partition sensitivity

The present rival audit detects overlap and incompatible relationships but does not yet make support invariant to how explanations are split or merged. Dividing one parent explanation into two mechanism variants can redistribute normalized mass even when the substantive evidence has not changed. Plan 043 therefore adds a hierarchical parent → mechanism-variant representation, parent-level reporting, and mandatory split/merge sensitivity.

3. Two models behind every observation

A likelihood judgment about evidence rests on more than a story about what caused the historical outcome. It also depends on how the alleged event became an observable trace. The method therefore distinguishes two models that are often collapsed:

Substantive causal model
How actors, conditions, decisions, and events produced later events and the outcome of interest.
Trace-production model
How an event or claim was solicited, recorded, copied, preserved, translated, retrieved, and extracted into the evidence available to the analysis.

This distinction matters because a trace can be present for at least two reasons: the underlying event occurred and was recorded, or the claim was produced even though the event did not occur as stated. Propaganda, bureaucratic formulae, retrospective rationalization, genre conventions, selective survival, and extraction error create false-positive channels. Conversely, censorship, secrecy, destruction, and archival absence can suppress traces of events that did occur.

P(observed trace | H) = P(fact | H) × P(trace | fact, H) + P(¬fact | H) × P(trace | ¬fact, H)

Retaining H in both trace channels is essential: even when the underlying fact is held fixed, different explanations can imply different incentives to fabricate, suppress, classify, standardize, or preserve its record. The current workbench captures source provenance, coverage, dependence lineage, and source-silence opportunity, but it does not yet represent this complete typed trace-production model. That is an explicit future-version target, not a capability implied by the present interface.

4. Evidence and experimental diagnostic scoring

Every evidentiary proposition must resolve to accepted source text and stable provenance. The exact quotation is a citation anchor, not the whole evidentiary object. Meaning can depend on the surrounding paragraph, genre, authorship, audience, purpose, competence, authenticity, temporal proximity, translation, editorial history, and incentives to misrepresent. The current exact-anchor requirement provides strong provenance but only partial contextual source criticism; Plan 043 makes contextual windows and those source fields mandatory for update-bearing evidence.

For each evidence item, the testing stage proposes one positive vector of relative evidentiary scores across all rivals. Pairwise ratios are derived from that common vector rather than elicited separately. This matters when there are three or more hypotheses: independent pairwise judgments can violate reciprocity or transitivity. Equal vector components mean the model judges the evidence nondiscriminating between those rivals.

experimental score ratioₘ(Hᵢ,Hⱼ) = qₘ,ᵢ / qₘ,ⱼ = exp(ηₘ,ᵢ − ηₘ,ⱼ), with qₘ,ᵢ > 0

This construction establishes algebraic coherence. It does not establish calibration: q may be a stable model score without being proportional to P(E | H). Paci's 2026 benchmark found that GPT judgments could be internally coherent and run-stable while remaining substantially different from expert estimates, predominantly overweighting evidence and sometimes changing the leading hypothesis. Automated magnitude estimation is therefore the principal validity bottleneck.

Priors remain conceptually distinct from evidence scores. Typed prior specifications retain their method, weights, rationale, and source provenance so the same source is not silently used twice. The current deterministic layer aggregates centered scores and exposes sensitivity:

experimental support(Hᵢ) ∝ prior weight(Hᵢ) × ∏ pooled score contributionₘ,ᵢ

The resulting weights are normalized comparative evidentiary support within the declared rival set. They are not posterior probabilities or posterior odds. Legacy internal fields still use likelihood/posterior names; those implementation names do not license the stronger public interpretation.

Diagnostic tests as a legible interpretation layer

Hoop, smoking-gun, doubly-decisive, and straw-in-the-wind tests summarize the shape of a likelihood comparison. They do not replace the likelihood model. A hoop test describes evidence strongly expected if a hypothesis is correct; failure is damaging, while passage may be weak support. A smoking-gun trace is unlikely under rivals; its presence is strongly supportive, while its absence may say little. The labels make diagnostic logic readable, while the multi-rival vector carries the coherent comparison.

Raw judgment and policy transformations

The system retains the originally elicited vector and records each deterministic transformation separately: relevance adjustment, dependence pooling, eligibility decisions, caps, and sensitivity perturbations. No workload cutoff or display convention is treated as a law of process tracing. A conclusion that changes under plausible priors, likelihood magnitudes, or guardrail settings is reported as fragile.

Dependence and repetition

A valid joint likelihood generally factors conditionally:

P(E₁,…,Eₘ | H) = ∏ₘ P(Eₘ | E₁,…,Eₘ₋₁,H)

The current system instead pools evidence through a claim → event → source → author hierarchy. This reduces obvious duplication, but a scalar discount is not equivalent to a joint likelihood model. Dependence may also be hypothesis-specific—for example, two documents may share one propaganda origin under one rival but represent separate observation channels under another. Plan 043 replaces the hierarchy as the sole representation with an evidentiary-ancestry graph, joint cluster judgments where needed, hypothesis-specific dependence, and leave-one-cluster-out sensitivity.

Absence of evidence

Source silence currently enters numerical aggregation only when an admitted source was positioned to record the predicted trace, the search is complete, no partial observation exists, and a separate audit accepts that opportunity-to-observe judgment. Missing source classes limit the conclusion; they do not prove absence in the world. A fuller trace-production model will replace the binary opportunity gate with rival-relative probabilities of nonappearance while preserving qualitative fallback when observability cannot be defended.

Selection and corpus conditionality

Exact quotations do not prevent selection on diagnosticity or latent model knowledge. A future version will retain the complete candidate-evidence universe—including rejected and nondiagnostic passages—freeze the retrieval and stopping rule, and audit inclusion independently of score extremity. It will also compare famous-name packets with anonymized entities and dates, fictional substitutions, permuted rival labels, and counterfactual packet variants. Material score changes caused by “Napoleon” rather than the admitted evidence violate the intended corpus-conditional claim.

5. Mechanism reconstruction

The method constructs a forward temporal graph from source-grounded stages. Each transition is classified as observed sequence, supported causation, contested causation, or unresolved. Temporal order alone never certifies causation. An unresolved link must identify the evidence that would clarify the causal break.

This distinction prevents a fluent narrative from hiding inferential gaps. A source may establish that event A preceded event B while leaving open whether A caused B, whether a third process produced both, or whether the relevant transmission is missing from the admitted record.

The graph is a qualitative causal audit, not a machine for reading likelihood magnitudes from topology. Without defensible parameters, a graph can reveal a missing pathway, confound, temporal contradiction, or translation loss; it cannot determine whether a likelihood ratio should be two or five. The narrative estimator supplies the magnitude, while the graph forces the causal and evidentiary structure behind that judgment into inspectable form.

The current category supported causation is still a corpus-conditional model judgment. Plan 043 will require each promoted edge to face a miniature rival-relative test: direct transmission versus a plausible alternative transmission, a common cause, and no causal connection. Until that control exists, the graph should be read as a typed and separately reviewed causal argument—not as independently identified causal structure.

6. Separate audit, traceability, and reproducibility

The component that produces a judgment does not certify it. Separate review stages inspect the rival partition, every materially discriminating likelihood comparison, the mechanism stages and edges, and the central claims in the final synthesis. Reviews are bound to the exact candidate artifacts and accepted source spans.

  1. Audit whether the rival set supports the proposed comparison.
  2. Review discriminating evidence against the exact quotation and opposed prediction.
  3. Review every mechanism stage and transition independently.
  4. Check terminal claims against their accepted evidence anchors before publication.

Rejected judgments, corrections, attempts, artifact hashes, model routes, costs, and final outputs remain inspectable. This creates traceability of one analytical execution; it does not by itself establish rerun reproducibility, because proprietary model versions, stochastic inference, retrieval indices, and tools can change. Future artifacts will bind exact model and version, prompts, reasoning and decoding settings, tool/runtime versions, source hashes, retrieval-index version, and seed where supported, then measure repeated-run variation.

Three meanings of auditability

Process auditability
Can a reviewer determine what stages ran, with which inputs, models, prompts, costs, and artifact versions?
Interpretive auditability
Can a reviewer trace a specific likelihood, mechanism link, or conclusion to the exact source text and contest its reasoning?
Interpretive validity
Are those judgments substantively defensible? Architecture can expose this question, but only representative empirical validation can answer it.

Independence is graded, not assumed

Separate prompts and blinded roles provide procedural separation by reducing anchoring and self-certification. Different model families can add model independence. Epistemic independence—different information or differently trained priors—is much harder: several models may reproduce the same historiographic stereotype. The default label is therefore separate audit. “Independent audit” is reserved for explicitly blinded cross-family or human review with recorded independence provenance. Agreement is supporting evidence, not proof, and same-model repetition is never described as independent corroboration.

7. Uncertainty, robustness, and falsifiability

Uncertainty enters through the source packet, rival partition, priors, diagnostic likelihoods, dependence assumptions, mechanism translation, and missingness judgments. The report keeps these sources separate rather than compressing them into one generic confidence score. It emphasizes whole-number comparative support, sensitivity ranges, rank stability, contested mechanism edges, and explicit source gaps.

A high comparative weight can coexist with weak evidentiary resolution when every rival is poorly observed or when one explanation merely loses less badly. Conversely, an indeterminate conclusion can be methodologically successful when the admitted record does not discriminate. Contract completeness, historical argument quality, and evidentiary resolution are separate evaluations.

Against ritualized Bayesianism

Equations, graphs, and audit logs can lend opaque judgments a false aura of objectivity. The architecture is answerable only if it can be tested against a simpler narrative-only baseline. Representative validation should measure calibration, discrimination, dependence detection, duplicate-overcounting control, rank accuracy, and expert contestability on held-out or adversarial cases. Thresholds and scoring rules must be fixed before evaluation; otherwise procedural sophistication becomes decoration.

The validation design has two strata. Synthetic and semi-synthetic cases expose a known data-generating process, allowing objective tests of calibration, rank recovery, duplication, dependence, trace production, and missingness. Real historical cases use blinded expert adjudication to evaluate defensibility, context use, error discovery, agreement, and useful disagreement—not an imaginary gold-standard historical posterior.

The decision gate

The benchmark must be able to reject numerical use. The post-MVP Plan 043 design predeclares tolerances for magnitude bias, rank error, proper-scoring performance, anonymization and label-permutation sensitivity, split/merge instability, dependent-cluster influence, contextual reversals, and rerun variation. Its corrected full benchmark remains unexecuted and is not an MVP release criterion. If future calibrated automation does not beat a simpler qualitative baseline on held-out evidence—or if these stress tests materially change the leading conclusion—the default product will suppress normalized numbers and retain scoring only as a structured deliberation aid.

What the unfamiliar-corpus pilot actually showed

On 17 August 2026, the hosted workbench received four previously uninstalled documents about Portugal's 25 April 1974 regime collapse. Eight accepted Gemini provider calls completed extraction through the experimental update for $0.110634; a ninth call stopped at mechanism reconstruction when the provider returned HTTP 503. The attempt was retained without a model switch or substantive retry, so it produced no final causal narrative.

The partial result is methodologically informative. Nine propositions from all four sources resolved to exact admitted quotations. Because the full corpus also generated the four candidate explanations, every proposition was excluded from diagnostic updating and all five displayed weights—including the residual—remained at their equal starting value. That equality is not a finding. The rival audit also classified all six pairs as potentially complementary and two pairs as absorptive, even while allowing the exploratory workflow to continue because each pair had an opposed prediction. This is direct implementation evidence that the present exploratory path can generate a plausible research agenda but does not yet implement the target typed relationships among compatible, conjunctural, nested, and sequential mechanisms.

8. What the result can—and cannot—establish

A completed run establishes what the admitted corpus supports, weakens, or leaves unresolved; how the declared rivals compare under the accepted diagnostic judgments; which causal links remain contested; and what evidence would most improve the inquiry.

Corpus conditionalThe finding is bounded by source coverage, provenance, and known gaps.
Experimental numerical layerCurrent support aggregates coherent model scores; it is not a calibrated likelihood or posterior probability.
Within-caseThe analysis does not estimate a population causal effect.
Argument, not oracleAutomation exposes judgments and uncertainty rather than eliminating them.

9. Extension to comparative studies

Process tracing remains a within-case method. Cross-case inference therefore begins only after each case has a separately inspectable full causal study and then applies a separate selection and comparison design. A compact trace is a useful intermediate projection, but it does not replace the rival specification, evidence testing, mechanism qualification, conclusion, and uncertainty analysis required of a completed case. A common discovery corpus may orient case selection, but it is not reused as independent case evidence.

discovery corpus → case source packets → full within-case studies → bounded cross-case synthesis

The ten-case revolution design

The current study asks: across bounded revolutionary and comparison episodes, what combinations and sequences distinguish episodes that establish a materially different governing order from episodes that do not? The unit is one connected challenge to an identified incumbent sovereign order with an observable mobilization, response, and endpoint. “Transformation” means that challengers displaced that order and established materially different governing authority within the frozen episode boundary. Mobilization, violence, leadership turnover, territorial control, or the label revolution is not sufficient.

The eligible universe was first assembled from three declared English-language discovery surfaces and reconciled into 1,963 provisional episodes. Twenty finalists then faced eligibility and pre-outcome comparability review before the separately retained endpoint ledger was joined. The final five–five balance is purposive contrast, not genuine outcome blinding, random sampling, prevalence, or a population estimate: names and boundary descriptions could reveal outcomes, and source availability was a feasibility constraint rather than a reason to select a famous case.

Selection follows structured, focused and theory-relative logic. Cases fill declared roles—anchor pathway, similar-context opposite outcome, same-polity temporal contrast, alternative path, external intervention, near miss, deviant non-transformation, and construct boundary—rather than pretending that one generic “most similar” rule solves every inferential problem. Replacements must fill the same role and preserve the frozen outcome and block design; inconvenient cases cannot be swapped merely for richer or easier sources.

Focused blockCasesInferential use and limit
Anchor opposite pathwayRomania 1989; Myanmar 1988Stable opposite-endpoint contrast in mobilization, coercion, elite alignment, and authority transfer; not an effect estimate.
Russian sequenceRussia 1905; February 1917Same-polity contrast in concession, coercive cohesion, organization, and displacement; shared institutions and historical sequence make the cases dependent.
Late Cold War contrastsRomania; Myanmar; Hungary 1956; East Germany 1953; Velvet RevolutionCompares domestic alignment, external coercion, information, restoration, and displacement while retaining bloc structure, diffusion, and historiographic dependence.
Transformation and boundary scopeFrench Revolution; Iranian Revolution; Velvet Revolution; May 1968Probes different transformation configurations. May 1968 is used only to test the construct boundary, not as substantive evidence for a general protest mechanism.
Non-transformation devianceMyanmar; Russia 1905; Hungary 1956; East Germany 1953Asks whether serious but restored challenges strain the same explanation in different ways; no case scores are averaged.

What crosses the case boundary

Within each completed case, rivals are tested against source evidence and the temporal mechanism is qualified. Across cases, the system compares only declared constructs, configurations, and sequences carried by those full studies. It does not pool case-level support weights, create one ten-case rival simplex, count overlapping blocks as independent observations, or infer necessity and sufficiency from recurrence. Every cross-case observation must step back to named cases, then to their exact evidence anchors; dependence and construct failures can exclude a case from a particular claim without removing its value as a boundary test.

The current implementation has ten completed source-limited full case studies and a separate full 18 Brumaire depth reference. Its five-block synthesis was rerun from those ten terminal case artifacts and separately audited. The result remains source-limited: it is a bounded comparative interpretation of the installed corpus, not a population estimate or a general theory of revolution.

When many cases are genuinely comparable, the project has a separate formal path built around explicit causal models and CausalQueries. That path can target identified counterfactual or population estimands only when common variables, case comparability, measurement, and model assumptions justify them. A single case is never sent through both engines and presented as if narrative support and a population causal effect were the same quantity.

10. Present implementation and methodological frontier

The public method distinguishes demonstrated controls from the quality-optimal architecture that guides future work. “Implemented” means the software path exists and has focused technical evidence; it does not mean that substantive historical validity has been established.

CapabilityStatusCurrent boundary
Inference-design and exposure lineageImplementedPublic output changes with theory-first, held-out, or exploratory status; exploratory discovery is not presented as independent confirmation.
Coherent multi-rival scoringImplementedOne positive score vector per item, deterministic log-space aggregation, typed priors, and sensitivity readout. Calibration as P(E | H) is not established.
Separate discriminator and mechanism auditsImplementedComplete typed review with exact-quote and artifact-hash binding; cross-family or epistemic independence is not guaranteed.
Dependence and source silencePartialHierarchical pooling and strictly gated source-silence logic exist; pooling is an approximation and dependence remains scalar rather than hypothesis-specific.
MVP prompt-behavior smokeDiagnostic onlyEight obvious behaviors appeared plausible, but the label-permutation retry changed a unique lead into an audited tie. Numerical support is therefore withheld from the MVP headline; retained scores remain inspectable and do not establish calibration.
Score uncertainty and calibrationPlan 043Point judgments, prior/driver sensitivity, and rank stability are exposed; held-out calibration, Paci-compatible comparison, bands, and joint propagation remain unfinished.
Context, selection, and contamination controlsPlan 043Exact anchors and source provenance exist; contextual packets, complete candidate custody, stopping rules, and anonymization/counterfactual probes are planned.
Partition, residual, and mechanism-edge sensitivityPlan 043Rival relationships and edge categories exist; hierarchical split/merge tests, a predictive residual, and local edge alternatives remain planned.
Trace-production modelPlan 043Provenance and coverage are retained, but causal production, recording, survival, hypothesis-specific incentives, and false-positive channels are not yet a complete typed model.
Representative validity benchmarkPlan 043Technical integrity and demonstrations exist; known-DGP calibration, blinded historical review, stress tests, and decision sign-off remain to be earned.

Read the method through an applied case

The 18 Brumaire report compares three predeclared explanations across 82 evidence records and retains its fragile lead, contested links, exact evidence, and next-source agenda.

References

  1. David Collier, “Understanding Process Tracing.”
  2. Andrew Bennett, “Process Tracing for Program Evaluation.”
  3. Derek Beach and Rasmus Brun Pedersen, Process-Tracing Methods.
  4. Sherry Zaks, “Relationships Among Rivals.”
  5. Sherry Zaks, “Updating Bayesian(s): A Critical Evaluation of Bayesian Process Tracing.”
  6. Tasha Fairfield and Andrew Charman, “Explicit Bayesian Analysis for Process Tracing.”
  7. Bennett, Charman, and Fairfield, “Understanding Bayesianism.”
  8. Simone Paci, “Is Research Safe in the AI Revolution? LLMs Fall Short of Expert Benchmarks in Scientific Evidence Evaluation.”
  9. Tasha Fairfield and Andrew Charman, Social Inquiry and Bayesian Inference.
  10. Macartan Humphreys and Alan Jacobs, “Mixing Methods: A Bayesian Approach.”
  11. Macartan Humphreys and Alan Jacobs, Integrated Inferences.
  12. John Gerring, “Case Selection for Case-Study Analysis: Qualitative and Quantitative Techniques.”
  13. Alexander George and Andrew Bennett, Case Studies and Theory Development in the Social Sciences.
  14. Tulia Falleti and James Mahoney, “The Comparative Sequential Method.”
  15. Stephen Van Evera, Guide to Methods for Students of Political Science.
  16. James Mahoney, “The Logic of Process Tracing Tests in the Social Sciences.”

Start a source-based causal inquiry

Define one causal question and supply documents for one bounded case. The analysis preserves the corpus, considers rival explanations, and keeps every conclusion connected to its source.

What the analysis does

  • Compares plausible causal explanations against the same evidence
  • Reconstructs a proposed sequence without equating sequence with causation
  • Shows the exact source anchor behind each material observation
  • Ends with unresolved questions and the most useful next evidence

What it does not claim

  • It does not decide whether your corpus is complete or representative
  • It does not infer source reliability, independence, or opportunity to observe at upload
  • Numerical comparisons are experimental evidentiary scores, not calibrated truth probabilities
  • The result remains conditional on your question, rivals, and admitted documents

Define the inquiry

These four decisions define the research design.

Before you run: admitted document text is sent to the configured model provider for analysis. Do not upload confidential, restricted, or personally sensitive material to this hosted demonstration.
Run preflight: this demonstration requests a maximum spend of $0.30 for this analysis (the current server-side cap is being checked). It may take several minutes. You can close the page and reopen the saved analysis; completed stages remain available.
Operator settings and repository-input compatibility
Technical stage reference
StageModuleInputOutputLLM?
1 · Extractpass_extract.pySource textEvidence items, actors, events, mechanisms, causal edgesYes
2 · Hypothesizepass_hypothesize.pyExtractionCompeting causal hypotheses with observable predictionsYes
2.5 · Partitionpass_partition.pyHypothesesAccepted rival set and opposed prediction contrastsYes
3 · Testpass_test.pyExtraction + hypothesesExperimental relative-score vector per evidence item; rejected discriminators trigger bounded re-elicitationYes
3a · Audit splitspass_discriminator_audit.pyExact quotes + proposed score differences + rival predictionsAccepted/rejected judgments and bounded re-elicitation historyYes
3b · Absencepass_absence.pyHypotheses + extractionMissing predicted evidence, source-opportunity judgment, and qualitative fallback when numerical admission failsYes
4 · Aggregate supportbayesian.pyEffective score matrix + priorsNormalized comparative evidentiary support, sensitivity bands, and robustness labelsNo — pure math
4.5 · Mechanismpass_mechanism.pyExtraction + rivals + accepted testingTemporally ordered stages and status-bearing causal transitionsYes
5 · Synthesizepass_synthesize.pyAll prior outputsWritten narrative, per-hypothesis verdicts, steelman casesYes
6 · Refinepass_refine.pySource text + first-pass resultsDelta (new evidence, reinterpretations) → re-runs passes 3–5Yes (optional)

Run progress

Follow the analysis from source extraction through reviewed publication. Open a stage for its status and result.

Create or resume a run, then select a stage to inspect its purpose and result.
1 · Extract
pass_extract.py · structured extraction calls
Complete
Input
Source text · 47,000 words
brumaire_18.txt · Source packet: 6 markers (A–F)
Research question: "Why did the 18 Brumaire coup succeed?"
Output
59 evidence items · 14 actors · 11 events · 4 mechanisms
12 causal edges · outcome: "Bonaparte consolidated power as First Consul"

What this does: Reads the bounded source corpus through typed extraction passes and retains historically observable items — evidence quotes, named actors and their roles, discrete events with dates, causal mechanisms, and directed causal edges. Each evidence item gets a relevance score and a diagnostic type (hoop / smoking gun / doubly decisive / straw in the wind).

Source fidelity rule: every evidence item must quote or closely paraphrase the input text. No hallucinated evidence. Circular evidence (derived from interpretation) is flagged.

★ Demo — showing Brumaire fixture data

Evidence sample (top 5 of 59)

IDSource textTypeRelevance
evi_murat_grenadiers19 Brumaire: Murat entered the Orangerie and announced that the Council was dissolvedSmoking gun0.92
evi_sieyes_directorMay 1799: Sieyès elected to Directory after years of republican constitutional draftingStraw0.88
evi_first_consul_powersDec 1799: First Consul promulgated laws; other consuls advisory onlyHoop0.95
evi_conspirators_rue_victoireEvening 17 Brumaire: Conspirators assembled at rue de la VictoireSmoking gun0.94
evi_ancient_regime_17891789: Estates-General convened; Third Estate declared National AssemblyStraw0.22 (below threshold)
2 · Hypothesize
pass_hypothesize.py · LLM call · ~1–2 min
Complete
Input
ExtractionResult · 59 items · outcome defined
Optional: --theories file (legitimacy vacuum framework injected)
Output
5 hypotheses · 3 observable predictions each
At least one agency hypothesis (named individuals). Opposite-prediction check passed.

What this does: Generates 3–6 competing causal hypotheses from the extraction. Each must specify a causal mechanism (not just an outcome description), name at least one observable prediction, and be genuinely rivalrous — if two hypotheses predict the same evidence in the same direction, they must merge.

Guardrails: anti-tautology, anti-circularity, a shared pinned outcome, distinct mechanisms, observable predictions, and system-recorded formulation exposure.

★ Demo — showing Brumaire fixture data

Hypotheses

IDLabelMechanism
h1Civilian revisionSieyès-led constitutional redesign; Bonaparte was the military guarantor Sieyès needed
h2Military conversionBonaparte's personalist military authority converted directly into political dominance via 18 Brumaire
h3Path closureContingent institutional failures progressively closed alternatives until coup was only viable exit
h4Legal engineeringElite legal-engineering by experienced jurists produced constitutional institutionalization of personal rule
h5Popular endorsementBroad popular demand for order and stability provided the legitimating base for Bonapartist rule
2.5 · Partition
pass_partition.py · rival-space validity gate
Complete
Input
Hypotheses + observable predictions
Output
Rival-pair audit + opposed prediction contrasts

Decision: Do the hypotheses form genuine rival causal configurations with concrete opposed predictions?

Demo · accepted rival-space fixture

Gate result

Accepted fixture partition with typed opposed predictions.

3 · Test
pass_test.py · bounded matrix chunks + corpus lineage · ~3–5 min
Complete
Input
59 evidence items + 5 hypotheses
Each bounded row chunk retains all hypotheses; exact evidence coverage is checked after recomposition.
Output
59 × 5 LR matrix · 6 below threshold
Relative likelihood per (evidence, hypothesis) pair. Pairwise LR cap = 20×.

What this does: For each evidence item, asks the LLM: "How likely is this evidence under each hypothesis, relative to the others?" Returns a likelihood vector, not independent pairwise scores — so ratios between hypotheses are coherent by construction.

Relevance gate: items with relevance < 0.4 are forced to LR = 1.0 (uninformative). Relevance = min(temporal, causal-domain). Cluster detection groups dependent evidence (e.g. all evidence from a single chain of events).

★ Demo — showing Brumaire fixture data
Cluster A detected: 8 Bonaparte military-career evidence items grouped as dependent (strength 0.70). These are pooled in the Bayesian update to avoid double-counting.

Likelihood matrix (top 8 rows by discriminating power)

Loading…

3a · Independent Discriminator Audit
pass_discriminator_audit.py · exact-quote semantic gate
Complete
Input
Every proposed pairwise split ≥2×
Output
Quote-level judgments + re-elicitation lineage

Decision: Does the exact quote make the favored account materially more likely than its named rival on the accepted opposed prediction?

Demo · no semantic audit fixture loaded

Exact-quote judgments

Run this stage to inspect candidates.
Evidence / quoteRival contrastProposed splitJudgmentReason
No live audit loaded.

Audit attempts

No live audit loaded.
3b · Absence
pass_absence.py · LLM call · ~1 min
Complete
Input
5 hypotheses + 59 evidence items
Each hypothesis's observable predictions checked against what was actually extracted
Output
3 failed hoop tests · qualitative severity ratings
Findings feed synthesis narrative only — NOT Bayesian updating (speculative LR assignment avoided)

What this does: Evaluates what predicted evidence is missing from the text. If h5 predicts popular demonstrations of support but the text contains none, that is a failed hoop test. Each absence finding is rated "damaging", "notable", or "minor" with a reasoning note about whether the text would contain this evidence if it existed.

Design constraint: absence findings are qualitative only. They inform the synthesis narrative but do not receive speculative LR values — which would add noise to the Bayesian update.

★ Demo — showing Brumaire fixture data

Absence findings

HypothesisMissing predictionSeverityReasoning
h5 · Popular endorsementPopular demonstrations of active support for BonaparteDamagingA text this detailed about the coup would mention mass rallies if they occurred; their absence is informative.
h1 · Civilian revisionSieyès drafting sessions with jurists post-coupNotablePost-coup constitution was drafted quickly; evidence of Sieyès-led sessions may exist in sources not covered.
h3 · Path closureDocumented attempts at non-coup constitutional reform in 1799NotableText focuses on coup itself; reform attempts might appear in broader political history sources.
4 · Bayesian Update
bayesian.py · pure math, no LLM · <1 sec
Complete
Input
LR matrix + priors · dependence clusters
Uniform priors (default). Dependence pooling applied to Cluster A. H0 residual included.
Output
Posterior distribution · sensitivity bands · robustness labels
Log-space softmax update. Order-invariant. Prior sensitivity: ±2× swing, rank stable.

What this does: Applies a coherent joint Bayesian update in log space: post_i = softmax(log prior_i + Σ log LR_i). Order-invariant. Per-evidence LRs are derived from the likelihood vectors as relative_likelihood / geomean(vector), so pairwise ratios are coherent by construction.

Then runs sensitivity analysis: for each hypothesis, perturbs its top 3 driver LRs ±50% and reports the posterior range. Robustness classification (robust / moderate / fragile) is mechanical — not LLM judgment.

★ Demo — showing Brumaire fixture data
Fragile dominant result. h2 leads at 0.994 but is fragile — driven by many weak LRs, not decisive smoking-gun tests. Under perturbation h1 can reach 0.44. Treat as a ranking, not a settled causal conclusion.

Comparative support (posteriors + sensitivity bands)

Shaded band = posterior range under ±50% perturbation of top LR drivers. Not a probability of truth — a comparative ranking.

4.5 · Temporal Mechanism DAG
pass_mechanism.py · typed LLM output · source-bounded edge assessment
Complete
Input
Accepted trace state · extraction · rivals · discriminator audit
Output
MechanismTraceResult · ordered stages · assessed arrows · missing tests

What this does: Converts the accepted evidence into a forward-only temporal DAG, then separately grades every transition as observed sequence, supported causal link, contested causal link, or unresolved link.

Reading rule: left-to-right order is necessary for causation, not sufficient. Evidence IDs ground what is observed; unresolved arrows name the next trace needed.

Demo · run a real case to inspect its typed mechanism trace

Mechanism stage lane

The live result renders here with a canonical edge table. Isolated nodes are invalid at this stage.

4.6 · Independent Mechanism Audit
pass_mechanism_audit.py · separate LLM role · hash-bound semantic gate
Complete
Input
Candidate DAG · exact corpus · frozen rivals · discriminator audit
Output
MechanismAuditResolution · final DAG · corrections · omissions · hashes

What this does: A separate semantic role checks every stage and transition against the accepted corpus. It may accept or weaken an edge, but it cannot strengthen one.

Failure rule: a material omitted stage or transition triggers bounded full-DAG replacement and re-audit, then blocks synthesis if unresolved.

Demo · run a real case to inspect edge judgments and correction lineage

Independent edge dispositions

The live resolution renders here as audit lineage plus stage, edge, correction, and omission tables. The raw typed contract remains available for inspection.

5 · Synthesize
pass_synthesize.py · LLM call · ~2–3 min
Complete
Input
All prior outputs · posteriors · absence findings
Output
Written narrative · per-hypothesis verdicts · steelmans

What this does: Writes the analytical narrative. The LLM has access to all evidence, hypotheses, LR vectors, posteriors, sensitivity ranges, and absence findings. It produces: a synthesis paragraph, a verdict for each hypothesis ("supported / partially supported / not supported / eliminated"), and a mandatory steelman case for every hypothesis — even eliminated ones.

Verdict calibration: verdict_calibration.py deterministically downgrades any verdict label that overstates the computed comparative support. The LLM writes reasoning; it does not have final authority on verdict labels.

★ Demo — showing Brumaire fixture data

Synthesis narrative (excerpt)

The evidence overwhelmingly supports h2 (Bonaparte's personalist military conversion) as the decisive causal mechanism of 18 Brumaire. The concentration of military command, the rue de la Victoire conspiracy, and Murat's use of grenadiers to dissolve the Council of Five Hundred all point to deliberate military agency rather than civilian constitutional revision. However, this result is classified fragile: it rests on a broad accumulation of weak-to-moderate LRs rather than a few decisive smoking-gun tests. The hypothesis cannot be considered settled absent higher-diagnosticity primary sources documenting Bonaparte's intent and planning.

h4 (Legal engineering) receives marginal support: the constitutional architecture of Year VIII does show elite legal craftsmanship, but this appears to be a consequence of the coup rather than its cause. h1 (Civilian revision) has a non-trivial sensitivity range (up to 0.44 under perturbation) and should not be dismissed without stronger discriminatory tests.

Verdicts

HypothesisVerdictPosteriorRobustness
h2 · Military conversionSupported (fragile)0.9942Fragile
h1 · Civilian revisionPartially supported0.0047Fragile
h4 · Legal engineeringNot supported0.0008Moderate
h3 · Path closureNot supported0.0003Fragile
h5 · Popular endorsementEliminated<0.0001Robust

Evidence provenance

6 · Refine
pass_refine.py · LLM call · ~5–7 min · optional (--refine flag)
Not run
Input
Source text + condensed first-pass results
~33K tokens. Second reading with full analytical context loaded.
Output
Delta · re-run of passes 3–5
New evidence (evi_ref_ prefix), reinterpretations, spurious removals, hypothesis refinements

What this does: Second reading of the source text with the full first-pass context loaded. The LLM produces a structured delta: new evidence items it missed on the first pass, reinterpretations of existing items, spurious items to remove, and hypothesis refinements. Then passes 3–5 re-run on the updated extraction.

Pre-refinement artifacts are saved to pre_refine/ before being overwritten, enabling the Delta Board in the workbench to show before/after posterior comparison.

Runs passes 3–5 again after delta is applied. ~$0.15 additional cost.
Refine stage not yet run for this session.
Run it above to see the delta board and before/after posteriors.
7 · Review & publish
central claim entailment review · LLM-backed publication gate
Not run
Input
Exact mechanism and synthesis claims
Every publishable span is checked against accepted source or deterministic artifact evidence.
Output
Accepted result.json + report.html
Publication fails loudly if any reviewed claim is overstated or unsupported.

What this does: inventories terminal claims, checks each atomic claim for source or artifact entailment, and publishes the final result only when the complete review is accepted.

When the persisted repair limit is one, a blocked review becomes exact constraints for one replacement mechanism, independent re-audit, and revised synthesis. A second block stops permanently; extraction, hypotheses, testing, and numerical support never change.

8 · Report
report.py · deterministic · <10 sec
Complete
Input
ProcessTracingResult · all passes merged
Full structured result including extraction, hypotheses, LRs, posteriors, synthesis, absence findings.
Output
result.json + report.html · temporal mechanism DAG
Temporal causal network, methodology-contract review, sensitivity table, and steelman cases.

What this does: Serialises the full ProcessTracingResult to result.json and renders a self-contained report.html. The report leads with the typed temporal mechanism DAG, then shows the evidence × hypothesis matrix, comparative support with sensitivity ranges, steelman cases, absence findings, provenance, source coverage, and hierarchical dependence.

No LLM calls — fully deterministic from the upstream structured results. Run again at any time to regenerate without re-running any LLM passes.

Report publication follows accepted terminal claim review.

result.json summary

FieldValue
run_idrun_20260624_200616_d4f6
evidence_items59 (6 below relevance threshold)
hypotheses5
dominant_hypothesish2 · Military conversion · posterior 0.9942 (fragile)
absence_findings3 (1 damaging, 2 notable)
dependence_clusters1 (Cluster A · 8 items · strength 0.70)
source_markers6 (A–F) · 49/59 evidence items traceable
historical_fixture_modelgemini/gemini-2.5-flash
historical_fixture_cost_usd~$0.08

report.html preview

[The generated report opens as a separate artifact. Its primary causal view is the deterministic temporal mechanism DAG; the force-layout extraction topology remains a secondary audit surface.]

One source anchor, two methodological uses

A real grounded-theory incident and a real process-tracing evidence item consumed through the same source-identity contract.

Candidate compatible Experimental consumer · no shared evidence score · no Data Contracts adoption

Loading the hash-bound comparison…

One challenge shape, two methodological consequences

See an authentic grounded-theory category revision beside process-tracing evidence that challenges a causal hypothesis.

Candidate compatible Experimental consumer · no method-independent scoring or inference
Critical restriction: Comparative likelihoods are not a shared challenge-strength score.

Loading the hash-bound comparison…

Evidence needed next

See what current evidence suggests, what remains uncertain, and what would test an explanation more convincingly.

Loading the P5 pre-run acquisition agenda…

Existing workflow

Post-result source evaluation

Requires a completed process-tracing result · ranked next-evidence needs

Acquisition Agenda

Not frozen
Freeze a completed result to build the agenda.

Candidates

0 retained
No retrieval attempts yet.

Admitted Evaluation Set

No admitted sources.

New-Evidence Result

Not run
The result compares uniform-prior support from newly admitted evidence with the frozen run for orientation.

Pipeline — Detailed Flow

Each edge labelled with the Pydantic type that crosses it. LLM / Math / IO badges per node.

← Click a stage to see its module, input type, and output type.

Schema — Pydantic Class Diagram

All data that flows between stages is typed. Every seam is a Pydantic model with Field(description=...) on every field.

Execution Sequence

Call order, LLM vs pure-math vs IO, and what crosses each boundary. Refine sub-sequence shown at bottom.

Confirm full analysis

This begins the remaining paid semantic stages. New calls stop visibly when recorded run cost reaches the aggregate admission cap.

Effective model
Reasoning policy
Published readout
Remaining stages
Call exposure
Repair limits
Admission cap

Call count is data-dependent because extraction, testing, claim review, retries, and accepted repairs may split into multiple calls. The shared client blocks new calls at the cap; an already in-flight provider call can settle above its remaining balance.

Review this research

Record what is convincing, what is not, and the exact change that would make the analysis more useful. Do not include confidential information or personal identifiers.