Config Agent can stage Brandon’s October 1 approval rules

This PR teaches the server-side Config Agent to see quote terms, attach runtime formulas to approval rules, and save staged rules instead of opening the editor.

PR #7535 Author @henry-dealops Branch henry/config-agent-brandon-fixes State open Files 50 Add / del +52,380 / -421 Area apps/server/src/dealops3

What it unlocks
The agent can now build approval rules scoped by quote-level terms like dealStructure and competitors, instead of inventing dead custom attributes.
What it hardens
Formula-backed rule conditions are resolved by stable formula key, validated like publish, and refused when the formula trigger does not match the rule subject.
Behavior change
“Stage it” now writes drafts through saveApprovalRules. The editor opens only when the user asks for it or gives an underspecified rule request.
Proof added
New evals replay Brandon’s Langchain approval request against a frozen Sep 28 template, plus quote-term cases and stricter grading on staged rules.
Quote terms
Runtime formulas
Approval rule writer
Agent loop
Eval harness

1. Why this exists

This is the third of four PRs split out of #7468, after roundNumber and the every-period check landed in #7533 / #7534.

Before

Brandon’s October 1 rules exposed three gaps

  • The agent saw Deal structure only as a SKU/product tag, not as the quote term reps choose.
  • It could not safely attach a custom runtime formula to an approval rule.
  • When asked to stage rules, it opened the editor about one third of the time, so nothing was saved.
After

The agent gets a saveable path

  • Quote terms are visible in the org snapshot with stored values and aliases.
  • Formula conditions name formulas by stable key and resolve to published definition IDs at save.
  • Failed writes trigger self-correction before the agent claims success.
The provided diff is truncated because the PR is large. This explainer is based on the PR description, file list, and visible diff; the largest addition is the frozen Langchain fixture JSON.

2. What changes

The production code is concentrated in Dealops 3 approval/config-agent server code; the 52k-line addition count is dominated by eval fixture data.

50 files changed
46k fixture JSON lines
7 new / updated eval cases
66/71 full Config Agent suite
Area
Old behavior
New behavior
Quote terms
Calculator deal terms and product-selection terms were missing from the snapshot.
Snapshot includes every quote-stored term, stored values, aliases, and SKU-shadowed keys.
Formulas
A formula key could be saved as a plain condition and never fire.
Formula conditions use stableKey, resolve to definitionId, and run publish-grade validation.
Rule staging
“Stage” sometimes opened the rule editor card instead of saving draft rules.
“Stage it” calls saveApprovalRules; editor opens only for explicit edit/open requests or underspecified requests.
Agent loop
A failed write could be followed by a confident final answer.
Failed write/validation tools force up to two self-correction rounds per turn, tracked per rule/formula/group.
Evals
Some cases graded plans rather than staged approval edges.
Brandon and image-matrix cases grade the draft rules actually added or changed.
Quote-term visibility
  • approvals/quoteTerms.ts
  • configAgent/orgSnapshot.ts
  • addCustomAttributeQuoteTerm.spec.ts
  • saveApprovalRulesQuoteTerms.spec.ts
Formula rule conditions
  • approvalFormulaConditions.ts
  • approvalFormulaSubjects.ts
  • approvalFormulaTools.ts
  • packages/types/v3/runtimeFormula.ts
Agent behavior
  • agents/toolFailures.ts
  • agents/agentResponse.ts
  • agents/orchestrator.ts
  • tools/modalTools.ts
Rule writer and evals
  • tools/approvalRuleLogic.ts
  • tools/approvalRuleTools.ts
  • evals/configAgent/harness/*
  • docs/quote-terms-next-steps.md

3. How it works

The new happy path is: discover quote terms, validate formula conditions, save draft rules, and retry if the model tries to paper over a failed tool call.

1 / Snapshot
Quote terms are listed
Product-selection terms and deal terms are keyed the way quote input stores them.
2 / Resolve
Formula keys become IDs
stableKey references are swapped for published formula definitionIds.
3 / Save
Rules are written as drafts
Rule conditions are recursively checked, including and, or, and not.
4 / Correct
Failures force another try
The orchestrator sends failed writes back to the model up to two times before final response.
Quote terms

approvals/quoteTerms.ts builds a canonical view of quote-stored fields:

  • product-selection terms;
  • calculator deal terms;
  • stored option values and aliases;
  • long option lists truncated at 25 with getQuoteTerm for the rest;
  • keys shadowed by SKU attributes in product scope.
Guardrails

Rule-writing tools now refuse configurations that would save but not work:

  • quote attributes that duplicate real quote terms;
  • quote-term conditions shadowed by product attributes in product scope;
  • exact-match labels that are not stored option values;
  • formula keys used as plain fields.
Formula condition shape

The agent names a formula by stable key, not by database ID:

{ "type": "formula",
  "formula": { "stableKey": "..." },
  "operator": "greaterThan",
  "value": 0.5 }

The tool resolves this to the published formula definition before save.

Subject matching

approvalFormulaSubjects.ts derives the rule trigger from compiled formula references.

  • input.currentProduct → product rule only.
  • input.currentPeriod → period rule only.
  • Mismatches are refused at save, not left for runtime surprise.
Self-correction contract

agents/toolFailures.ts normalizes failed tool results, including thrown tools, so telemetry and the runtime agree on what failed.

  • Failures are tracked per rule, formula, or group.
  • A retried rule can match by ID, description, or routing.
  • A successful save for one rule does not hide a refusal for another.
Last-round behavior

On the final model round, the agent does not retry again.

It returns the answer with a note naming what was not saved, instead of claiming the whole request succeeded.

4. Eval and fixture coverage

The PR adds a local-only historical baseline plus model evals that grade the staged artifacts rather than just the agent’s plan.

Frozen Langchain template

brandon-approval-sep28.json reconstructs the Sep 28 Langchain approval environment: specs, groups, formulas, pricebooks, catalog, and preview input.

46,441 added JSON lines are fixture state, not runtime logic.

Brandon cases
  • replay-langchain-oct1-no-reuse-001
  • langchain-oct1-build-with-new-formulas-001
  • replay-langchain-quote-deal-structure-001
Quote-term cases

approval-rule-calculator-term-001 checks that the agent stages a rule on the existing competitors term using stored value sierra.

Harness additions

approvalSandbox, beforeTurn: publishApprovalFormulaDrafts, ifPlanned: approve, and changedOnly make multi-turn approval evals grade the right state.

Run / suite Reported result Reviewer read
Full Config Agent suite on branch 1352a79 66/71 pass, 12 skipped Four failures match known-bad main cases; one image-matrix flake passed on two reruns.
Five new Brandon / quote-term cases All pass on branch They validate staged rules and routing targets, not approval simulation quote outcomes.
Jest 112 suites, 1,524 passing One catalog suite timeout under parallel load passes alone.
Mocha 14,099 passing One local call-context backfill failure in an unrelated area.
Typecheck / ESLint / e2e Clean typechecks, 0 ESLint errors, 14 e2e passing before latest merge Brandon baseline e2e is local-only and skipped in CI.

5. What it doesn’t change

Explicit non-goals
  • No approval simulation or test-run card. That is the next PR in the stack.
  • No client UI route, tRPC route, Prisma migration, or production database schema change is introduced here.
  • No product, list price, pricebook, or existing approval rule is changed by runtime code; the large catalog/spec data is eval fixture state.
  • No automatic publish of approval formulas or approval rules. Formula-backed rules still require published formulas.
  • No required id on every rule created by saveApprovalRules; that tool-contract change is a follow-up.
  • No approval-conflict detector replacement in this PR. detectConflicts is removed and a scoped version is called out as follow-up.

6. Risks / rollback / open questions

Behavior now affects every Config Agent turn
Self-correction, more forgiving reply parsing, and the false-success guard tweak run globally. This is intentional, but reviewers should look for latency, over-retry, and changed final-answer tone.
Risk

Retry matching is not exact yet.

Until rule IDs are required, a retried rule can clear a failure by matching ID, description, or routing. Duplicate rules with identical descriptions/routing are the edge case.

Risk

Editor behavior changes for users.

Users who relied on “stage” opening the editor will now get saved drafts unless they explicitly ask to open or edit the rule card.

Rollback

Reverting this PR restores prior staging/editor behavior and removes the new eval fixture and guardrails.

There is no schema migration to unwind.

Open follow-up

Require IDs on every saveApprovalRules create, add scoped conflict detection, and land approval simulation in the next PR.