Working paper · Reasoning in agentic systems

Codifying implicit decision reasoning: the tradeoffs nobody wrote down

Most codification captures what an expert can write down — steps, rules, checklists. But transcripts and memos record what was done and what was said about it, never the tradeoff that actually decided it. This lays out an approach for recovering that implicit reasoning — the criteria, their conditional weights, and the constraints — so an agent can reason in new situations instead of imitating old ones.

Claim 1 A behavior twin copies the policy. It reproduces what was done in situations that look familiar. It has nothing to say when the situation is new.
Claim 2 A reasoning twin recovers the reward. Inverse RL's move, applied to knowledge engineering: infer the criteria and their conditional weights, then re-derive the decision anywhere.
Claim 3 The twin is a state, not a snapshot. Each round of codification is a noisy observation of a hidden decision function that drifts. Shrink toward what you knew when the new evidence is noisy; adapt when it is precise.
Claim 4 Determinism is measured, not designed. Which parts become hard rules and which stay judgment falls out of the fitted, filtered model — it is not a separate design choice.
01 · Why this matters

Agents are being built from the wrong artifact

The dominant way to make an agent act like a specific expert is to feed it that expert's trail: chat transcripts, meeting recordings, written memos, a curated memory of past cases. A language model summarizes the trail into rules, procedures, and examples. The agent then follows them.

This works until the situation is one the trail never covered. Then the agent either freezes, applies the nearest rule out of scope, or invents a rule that sounds right. The failure is structural, not a matter of more data: the trail records outputs of reasoning and a partial narrative of it. The reasoning itself — a judgment over competing criteria under constraints — was never written down, because the expert never had to write it down, and often could not.

Behavior twin looks up the nearest past case; reasoning twin scores the new case against recovered criteria BEHAVIOR TWIN · COPIES THE POLICY New situation Deal D Nearest past case ≈ Deal B Copy action invest ✗ Wrong call. The ownership floor never appeared in any past case's narrative, so there is nothing to match on. REASONING TWIN · RECOVERS THE REWARD New situation Deal D Score criteria g_i(s, a) Weigh, constrain w(s) · h(s) ✓ Pass — and it can say why. Ownership 9% violates the recovered constraint h_own : ownership ≥ 15% ∧ lead
The same new deal, two kinds of twin. Matching on past cases has nothing to say about a criterion that was never contested before. Scoring against recovered criteria does.

Three findings from outside AI say this is not a fringe case. Polanyi: experts know more than they can tell. Argyris and Schön: what practitioners say governs their decisions and what actually governs them diverge reliably. Nisbett and Wilson: people confidently explain choices whose real causes they had no access to. A stated rationale is evidence, not ground truth.

The rest of this paper works through a single venture-capital decision to show where the hidden reasoning lives, then lays out a process — with a statistical model at its center and a language model doing the semantic lifting — for recovering it.

02 · A Series A decision

One partner, three deals, one slot

A general partner at a $200M early-stage fund has capacity for one more Series A lead this quarter. Three opportunities are on the table. The partner picks one, writes the memo, and the investment committee approves. Everything below is fictional but representative.

Deal AVertical SaaS · logistics
ARR
$4.0M
Growth (YoY)
2.5×
Post-money
$60M
Ownership at $8M
13%
Competing term sheets
3
Founder history
First-time
Passed
Deal BAI infrastructure
ARR
Pre-revenue
Growth (YoY)
Post-money
$25M
Ownership at $5M
20%
Competing term sheets
0
Founder history
Ex-CTO, prior portfolio exit
Invested — lead
Deal CFintech marketplace
ARR
$1.5M
Growth (YoY)
4.0×
Post-money
$30M
Ownership at $6M
20%
Competing term sheets
0
Founder history
Second-time, small exit
Passed

What the record shows

This is the material a codification pipeline would ingest: the investment memo and the partner meeting transcript. Both are honest. Both are incomplete.

Investment memo · excerpt
We recommend leading Deal B's Series A. The founder's prior execution as CTO of a portfolio company gives us unusual conviction in the team. The market timing for inference infrastructure is favorable and the entry price is attractive relative to comparable rounds. Deal A is a strong company but the round is competitive and priced accordingly. Deal C shows excellent growth but carries regulatory exposure we are not positioned to underwrite.
Partner meeting · transcript excerpt
GP: Honestly A is the "safest" business of the three. But we'd be one of four names on the cap table at 13%. I don't love that.
Partner 2: C's growth is real though.
GP: It is. I just… after last time, I'm not doing another money-movement business without the license in hand. Not this fund.
Partner 2: And B?
GP: I've worked with him. I know how he ships. And frankly we need an AI infra name in this vintage before we're out raising Fund IV.

What actually decided it

Read the transcript again. Four things are doing the real work, and none of them appear in the memo as criteria. Two are barely spoken at all.

Constraint
Ownership floor — roughly 15%, and the fund must lead."One of four names at 13% — I don't love that." Never written as a rule. Never violated in six years of this partner's deals. A near-hard constraint that reads in the transcript as a mild preference.
Constraint · conditional
No unlicensed money-movement risk — in this fund."After last time… not this fund." A constraint that did not exist two years ago, created by a prior loss, and scoped to the current vehicle. The memo calls it "regulatory exposure we are not positioned to underwrite," which sounds like analysis. It is a scar.
Weight · situational
Fund narrative weight — high because it is year 4 of a 5-year deployment period."We need an AI infra name in this vintage before we're out raising Fund IV." This criterion has near-zero weight in year 1 of a fund and dominant weight in year 4. The memo folds it into "market timing."
Weight · relational
Diligence confidence — a known founder collapses perceived execution risk."I know how he ships." The memo says "unusual conviction in the team," which is accurate but not operational: it does not say that prior working relationship substitutes for the two-week technical diligence the fund would otherwise require.
What actually drove the decisionIn the memoIn the transcriptOperating
Ownership floor ≥ 15%, must lead
No unlicensed money movement (this fund)
Fund-narrative weight, high in year 4
Known founder substitutes for diligence
Growth, price, team quality (the stated criteria)
absent veiled — present as a feeling or a euphemism stated actually decisive
The stated criteria are fully documented and only half-decisive. The decisive criteria are, at best, veiled. A pipeline that reads the left two columns codifies the wrong row.
The test that breaks the copy

Next quarter, Deal D arrives: AI infrastructure, founder is a known ex-CTO from the network, strong market narrative — and a competitive round at $80M post, where the fund could get 9% as a follower.

A behavior twin trained on the memo and transcript has learned "AI infra + known founder + market timing → invest." It recommends Deal D. The partner would pass without a second thought: the ownership floor is a constraint, and it is not in the artifact because it was never at risk in Deal B and therefore never discussed as a criterion.

03 · The codification gap

Why "what" and "why" are not enough

A decision has three layers. Today's pipelines read the top two and produce an artifact that looks like the third.

Three layers of a decision and which ones the record captures L1 · ACTION Invested in B; passed on A and C. Wire sent. Board seat taken. L2 · NARRATIVE "Team conviction, market timing, attractive price, regulatory exposure." L3 · DECISION FUNCTION Ownership ≥ 15% (constraint) · no unlicensed money movement (constraint, this fund) w(narrative | fund year 4) ≫ w(narrative | year 1) · known founder → skip tech diligence IN THE RECORD Fully — logs, wires, cap table IN THE RECORD Partially — and post-hoc IN THE RECORD No. Must be inferred. This is what the agent needs.
Codifying from transcripts reads L1 and L2. The agent has to act from L3. The distance between them is the codification gap.

The memo is a projection of L3 onto a document genre. It is written to persuade a committee, so it foregrounds criteria that sound like analysis (market timing) and suppresses ones that sound like bias (I know this founder) or institutional need (we need a logo for the next raise). This is not dishonesty. It is what memos are for.

The transcript is closer, but still only surfaces criteria that were contested in the room. The ownership floor was mentioned because Deal A threatened it. Had all three deals offered 20%, it would never have come up — and the codified artifact would have no idea it existed.

Four kinds of thing are systematically missing, and each is missing for a different reason:

Hidden elementDeal B exampleWhy the record misses it
Hard constraintsOwnership ≥ 15%; lead positionConstraints are only spoken when threatened. Most decisions never threaten them.
Conditional weightsFund-narrative weight depends on fund yearThe condition (year 4) is ambient context the speaker doesn't restate. The record has the weight at one point, never the function.
Tradeoff orderingNarrative > price > growth, this quarterOrderings are revealed by choices, not stated. One decision reveals one comparison.
SubstitutionsKnown founder substitutes for technical diligenceExpressed as a feeling ("I know how he ships"), not as the procedural exception it actually is.
Narrative weight as a step function of fund year; the memo captured one point on it R1 · YEAR ≤ 2 R2 · YEAR 3 R3 · YEAR ≥ 4 Y1 Y2 Y3 Y4 Y5 FUND YEAR (SITUATION) 0 0.5 1 w_narrative(s) What the agent needs: the whole function What the memo captured: one point, with the condition left implicit
"Conditional preference": the weight is a function of the situation, piecewise constant over regions. A record of one decision fixes one point on that function. The regions R1–R3 are themselves something to discover — the memo never names "fund year" as a variable.

This is the case for a different target. Inverse reinforcement learning made the same move two decades ago: don't imitate the demonstrator's policy, recover the reward function that makes the policy optimal, then re-derive the policy anywhere. A digital twin that copies behavior is apprenticeship without the reward. A reasoning twin recovers the reward — criteria, conditional weights, constraints — and derives the behavior. It transfers to Deal D. And it is arguable: a partner can look at "ownership floor 15%" and say "actually it's 12% for AI infra," which they cannot do with a policy.

04 · Improving the codification process

Recovering the decision function

Two tiers. Plain language explains what the process does and who does what. Full detail gives the model, the fitting procedure, and the math. Same content, different altitude.

The process has three roles, and the design principle is that each does only what it is good at.

LLM · semantics Statistical model · weighing Human · judgment

The language model reads the record and proposes what the criteria might be, what alternatives were on the table, and what question to ask next. It is good at meaning and bad at weighing. The statistical model takes the proposed criteria and the observed choices and estimates how much each criterion mattered, under which conditions. It is good at weighing and cannot invent a criterion. The human is never asked to state a weight — only to make a judgment on a concrete case. Experts are reliable at the second and unreliable at the first.

The seven-step codification loop 1 · LLM Build decision instances situation, choice, alternatives 2 · LLM Propose criteria over-propose; let the fit prune 3 · MODEL Fit conditional weights sparse, region-indexed 4 · MODEL + LLM Find conflicts → propose hidden context 5 · HUMAN Answer contrastive query "would you still, if…?" 6 · MODEL Place on determinism spectrum rule · judgment · constraint · escalate 7 · AGENT Deploy and monitor escalations, overrides NEW INSTANCES → REFIT ANSWER → REFIT
The loop never terminates. Every escalation and override at runtime is a new instance; the model is refit; the artifact is re-issued.
1
LLM · Build decision instances

Turn the record into a list of decisions, each with the situation, the choice, and — critically — the alternatives that were rejected. The transcript shows what was done, not what was passed on. The LLM reconstructs the alternatives from the record and, where the record is silent, proposes plausible ones for the human to confirm.

Deal B: situation = {fund year 4, three deals, one slot}; chosen = B; alternatives = {A, C, pass on all}. Plus the prior six years of this partner's deals, each reconstructed the same way.

2
LLM · Propose criteria and context features

Read the memos and transcripts and propose what the partner might be weighing — generously. Over-proposing is fine; the fit in step 3 will zero out what does not matter. Also propose situation features that might change the weights: fund year, sector heat, whether the partner has a personal relationship with the founder.

Proposed criteria: growth, price, ownership, team quality, founder relationship, market narrative, regulatory exposure, competitiveness of round. Proposed context: fund year, dry powder remaining, recent portfolio loss.

3
Statistical model · Fit conditional weights

Given the instances and the criteria, estimate how strongly each criterion pulled the choice, and whether that strength depends on the situation. The model is deliberately simple and interpretable — an expert produces dozens of decisions, not millions — and it outputs a small table: this criterion, this weight, under these conditions.

Output: ownership has a sharp cliff near 15% (looks like a constraint, not a preference). Market-narrative weight is near zero in fund years 1–2, dominant in years 4–5. Regulatory exposure flips from mild negative to near-veto after the loss in year 2.

4
Model + LLM · Find conflicts, propose hidden context

Look for pairs of decisions that look the same on the current features but went opposite ways. These are conflicts. Instead of averaging them away, ask the LLM: what could distinguish these two cases? Add its proposal as a feature and refit. A conflict is a missing variable until proven otherwise.

Two AI deals in year 1 with strong founders — both passed. Deal B in year 4 — invested. The features cannot separate them. LLM proposes "years remaining in deployment period." Refit: separation is clean.

Three AI deals overlap on two features; adding fund year separates the passes from the investment BEFORE · CURRENT FEATURES Same features, opposite choices — the fit cannot separate them FOUNDER STRENGTH → NARRATIVE FIT ↑ 2 passed · 1 invested AFTER · ADD "FUND YEAR" (LLM-PROPOSED) Clean separation — the conflict was a missing variable Y1Y2Y3Y4Y5 FUND YEAR → REGION SPLIT passed · w_narrative ≈ 0 invested · w_narrative ≈ 0.9
Step 4 in one picture. The residual cluster is not noise to average over; it is the strongest available evidence that a variable is missing. The LLM's job is to name it; the fit's job is to test it.
5
Human · Answer one contrastive question

Pick the single question whose answer would most reduce uncertainty in the weights — usually a minimal variation of a real case, at a suspected boundary — and ask the expert to judge it. Never "what are your criteria?" Always "would you still, if this one thing were different?"

"You led Deal B at 20%. Same deal, same founder, but $50M post and you get 10% as a follower. Still in?" — "No." One answer converts a suspected constraint into a confirmed one. "What if 14%?" — "…probably. If we lead." Threshold found.

Contrastive queries on ownership locate the constraint threshold by bisection CONTRASTIVE QUERIES ON ONE FEATURE · OWNERSHIP AT ENTRY τ ≈ 14–15% 5% 10% 15% 20% 25% Deal B as done · invested Q1 · "10%, follower?" — No Q3 · "12%?" — No ? Q2 · "14%, we lead?" — "…probably"
Step 5 in one picture. Three judgments on concrete variants of a real deal — never "what is your threshold?" — pin a constraint the partner never stated to a one-point band. Each answer is one row in the fit, with very low noise.
6
Statistical model · Place on the determinism spectrum

For each decision region, measure how predictable the choice is once the criteria are known. Highly predictable and stable → codify as a rule. Unpredictable even after context discovery → codify as guided judgment with the weights and contrasting examples attached. Irreversible downside → constraint, regardless. Section 06 covers this; section 05 makes it track the evidence over time.

7
Agent · Deploy, monitor, loop

The agent carries the criteria explicitly and scores new situations against them. Every time it escalates (situation outside the evidence) or is overridden (human disagrees), that is a new instance. Return to step 3.

Deal D arrives. Ownership at 9% violates the constraint. The agent passes and can say exactly why. No memo ever said "ownership floor."

Fitted weights · fund year 1
Growth
0.78
Price / ownership
0.64
Team quality
0.70
Founder relationship
0.30
Market narrative
0.08
Fitted weights · fund year 4
Growth
0.44
Price / ownership
0.52
Team quality
0.68
Founder relationship
0.62
Market narrative
0.91
Illustrative fitted weights for the same partner in two situations. The criteria did not change; the weights did. A behavior twin sees two inconsistent partners. A reasoning twin sees one partner and one variable — fund year — that was never in the memo. Ownership floor and regulatory exposure are constraints, not weights, and are not shown.
05 · The update machine

The decision function is a hidden state, and codification is a noisy reading of it

Section 04 recovers the decision function from one batch of evidence. But the partner's weights drift — a fund ages, a loss leaves a scar, a market heats up — and each round of codification is itself noisy: a few new deals, a handful of answers. Two mistakes are available: overwrite what you knew with a noisy new fit, or ignore real change because it disagrees with the old artifact. A state-space filter is the principled middle.

Think of the partner's real decision function as a hidden state — the weights and thresholds that are actually in force this quarter. You never see it directly. What you see, each period, is a reading of it: the weights the fit in section 04 produces from that period's deals and elicitation answers, together with how uncertain that fit is (the ± column in the fitted table).

The update rule is the one every navigator and every Kalman filter uses. Carry forward what you believed last period, allowing for some drift. Take the new reading. Then move your belief toward the reading by an amount that depends on which is more trustworthy: if the reading is noisy and your prior is sharp, barely move — shrink toward the prior. If the prior is vague and the reading is precise, move most of the way — adapt to the observation. The amount you move is the gain, and it is computed, not chosen.

Predict from the prior state, observe a new codification, compute the gain, update the state, carry it forward PRIOR STATE What we believed last period, plus allowance for drift x̂ₜ₋₁ , P + Q OBSERVATION This period's fit from §04, with its own uncertainty yₜ , R (the ± column) GAIN · THE DIAL How far to move: prior uncertainty vs. reading noise K = P / (P + R) UPDATED STATE Prior, nudged toward the reading by K x̂ₜ = x̂ + K (yₜ − x̂) CARRIED INTO NEXT PERIOD · t → t+1
One cycle of the update machine. Section 04 produces the observation; this loop decides how much of it to believe. The observation's noise R comes free from the fit's own posterior — nothing new has to be estimated.

The four cases

The gain reduces to a two-by-two: how sure were we before, and how clean is the new reading?

New reading is noisy
(few deals, no elicitation — R large)
New reading is precise
(elicitation answers, many deals — R small)
Prior is sharp
(P small)
K ≈ 0 · Shrink toward priorKeep what you knew. Log the anomaly.One odd deal in a quiet quarter does not move a weight that 20 prior deals established. It is recorded as a possible early signal, nothing more.
K mid · Blend — and testMove partway. If the gap is large, suspect a regime change.Two sharp beliefs that disagree is the interesting case. A large, repeated surprise means the drift allowance was too small — the world changed. Inflate it and ask the LLM to name what changed.
Prior is vague
(P large)
K mid · Blend weaklyNobody knows much. Ask a question.This is where the next contrastive query should go: whichever weight has both a vague prior and a noisy reading has the highest expected information gain.
K ≈ 1 · Adapt to observationTake the reading. This is what elicitation is for.A precise answer on a weight you were unsure about should move the state almost all the way. The regulatory scar after the year-2 loss lands here: one "not this fund" answer, low noise, vague prior.
Narrative weight over five fund years: noisy observations, filtered state, and a regime shift SURPRISE PERSISTS → DRIFT ALLOWANCE INFLATED Y1Y2Y3Y4Y5 00.51 noisy reading (one deal) → shrunk toward prior, K ≈ 0.1 precise reading (elicited) → state adapts, K ≈ 0.8 state uncertainty P filtered estimate x̂ (the twin) each period's codified reading y ± R
w_narrative for one partner across a fund's life, illustrative. Early readings are noisy and the state barely moves. In year 3 the readings start disagreeing with the prior, and keep disagreeing — that is a regime change, not noise, and the filter widens its drift allowance to follow. The year-4 elicitation answer is precise and the state snaps to it. The behavior twin would have flip-flopped on every reading; a frozen artifact would still say 0.08.

What this changes in the process

  • "Supersede" stops being a rule and becomes an update. Newer evidence overrides older evidence exactly in proportion to how precise it is. A single noisy quarter cannot erase a well-established weight; a single precise elicitation answer can.
  • Constraints move only on precise readings. A threshold like the 15% floor is part of the state, but with a very sharp prior. Instances alone will never shift it; only a low-noise elicitation answer will. That is the right behavior for something that functions as a rule.
  • Regime change gets detected, not discovered by accident. Persistent surprise is the temporal version of the conflict in step 4: a variable changed that the model does not track. The same LLM step names it — "post-loss", "year ≥ 4", "Fund IV raise underway".
  • Where to ask next is computed. The weight with the widest state uncertainty is the one whose next contrastive query pays most.
  • A partner's twin can borrow from the firm. With few deals, an individual's weights shrink toward the firm-wide prior; as their own evidence accumulates, they earn their own. Same gain, applied across people instead of across time.
06 · Deterministic vs. non-deterministic artifacts

Rigid where rigidity is safe, judgment where it must be

The last question in any codification is "should this be a rule the agent follows, or a judgment the agent makes?" With a fitted choice model, this stops being a design meeting and becomes a measurement.

The determinism spectrum from hard constraint to human escalation DETERMINISTIC NON-DETERMINISTIC Constraint"never X" Procedure"do A, then B" Scored heuristic"score > θ → X" Guided judgmentweights + examples Open judgmentagent decides Escalatehuman decides H ≈ 0 · irreversible H ≈ 0 · stable H low · 1 feature H high · reducible H high · irreducible out of provenance
H is the entropy of the fitted choice probability within a region. Position on the spectrum is determined by entropy, stability across refits, reversibility of the downside, and whether the situation is inside the evidence the model was fit on.

The placement rules

For each region R_k — each named situation the model found — ask five questions in order. The first that applies wins. Every input to these questions is a quantity the fitted, filtered model already produces.

QuestionIf yesWhy
Is the downside irreversible or high-blast-radius?Constraint on the downside, regardless of what else the region doesA wrong judgment call that can be undone is a learning event. One that cannot is a loss. Cost asymmetry beats predictability.
Is the choice near-deterministic (low entropy) and stable across refits?Procedure or scored heuristicThe expert always does this here. Judgment adds variance without adding value, and rules are cheap to run at volume.
Is the entropy high but the region reducible — a variable might still be found?Guided judgment: emit weights, the contrasting examples, and the open questionThe agent should decide, but with the tradeoff made explicit and the evidence attached, so its choice can be audited and its exceptions harvested.
Has the filter flagged a regime change on this region?Escalate and re-elicit; hold the prior artifact as provisionalPersistent surprise (section 05) means a variable changed that the model does not name yet. Acting on the old rule is acting on a stale state.
Is the situation outside the provenance set?Escalate — at runtime, dynamicallyAny region can be pushed to escalation for a specific case. A reasoning twin that knows what it does not know is worth more than one that guesses.

The VC decision, placed

ConstraintOwnership ≥ 15% and leadH ≈ 0 across 17 instances; downside (a passive minority position) is irreversible for the fund's model.
Constraint · scopedNo unlicensed money movementPost-loss region only. Confirmed by elicitation. Would have been a preference before year 2.
ProcedureDiligence checklistTwo-week technical diligence — with one codified exception: known founder from the network substitutes for it.
Guided judgmentTeam vs. price vs. growthHigh entropy, reducible. The agent scores with fitted weights for the current fund year and cites the closest prior deals.
EscalateNarrative bets in year 4+Partners disagree uniformly on how much fund story should weigh. Values conflict, not a missing variable. Investment committee decides.

Two failure modes bracket this. Over-codification produces an agent that applies a rule outside its provenance and cannot tell it has done so. Under-codification produces an agent that re-litigates settled questions inconsistently. The spectrum is the mechanism for placing each region between them deliberately — and, because placement is a function of the filtered state, re-placing it automatically each period as the evidence changes. Nobody re-decides "is this a rule?" in a meeting; the state does.

07 · Conclusion and what lies beyond

Codification is a living contract, not a document

The argument in one line: an agent built from transcripts inherits a policy; an agent built from a recovered reward inherits a way of deciding. The first is a behavior twin and fails at Deal D. The second is a reasoning twin and passes on Deal D for the right reason.

The single model — a conditional choice function with explicit constraints — unifies four things that today are handled separately and by hand: tradeoff weights (fitted), context (regions), conflict (residuals that point at missing variables), and the rule-versus-judgment decision (entropy). Wrapping it in a state-space filter adds the fifth: time. Each round of codification becomes a noisy reading of a drifting hidden state; the gain decides how much to believe it; persistent surprise flags a regime change and sends the LLM to name it. Provenance comes free, which means every element of the codified artifact can be disputed by the expert it claims to represent — and every element carries a date and an uncertainty.

What to look at next

The practical shift is small in mechanics and large in posture. Stop asking experts to write down their reasoning. Ask them to make judgments on cases the model chose, let the model do the weighing, and let the language model do the reading. The reasoning was never in the transcript. It was in the choices, waiting to be asked the right question.