Config Agent tool surface consolidation

PR #7387 cuts the CPQ Config Agent’s normal tool surface from 60 tools to 30 by injecting the context it used to fetch and replacing near-duplicate tools with consolidated read/write APIs.

Author: @henry-dealops PR: dealops#7387 Branch: henry/config-agent-tool-path-evals Area: apps/server/src/dealops3/configAgent Files: 63 State: open

What it changes

The agent binds fewer tools, starts with richer org and file context, and routes common mutations through consolidated APIs such as saveApprovalRules, saveApprovalGroup, and bulkStageCatalogProducts.

Why it matters

Every normal turn previously paid for roughly 17k tokens of tool definitions and often spent round 1 re-reading data already knowable before the LLM call.

Result

Normal turns drop from 60 tools to 30, tool definition tokens drop to about 11.6k, and full-suite pass@1 remains green with no introduced eval regressions.

Review focus

Review the semantic compatibility of the consolidated tools: patch behavior, approval rule cloning, live group names, file key reads, and snapshot fallback paths.

Orchestrator
Org snapshot
Conversation files
Consolidated tools
Eval harness
E2E tests

1. Why this exists

Before
  • Orchestrator bound 60 tools on a normal turn.
  • Tool definitions cost about 17.0k tokens every round.
  • Several tools overlapped by entity and operation.
  • Round 1 often fetched approval groups, rules, files, or config that could be provided up front.
After
  • Consolidated tools expose one read/write shape per kind of change.
  • Org snapshot includes approval and composite-price context.
  • Conversation files are listed in prompt context with stable keys.
  • Eval harness now measures the tool path, not just final outcome.
Diff note: the supplied diff is truncated. This explainer uses the PR description, file list, and visible diff chunks; it does not claim line-by-line coverage of omitted hunks.

2. What changes

Normal turn tools
60 → 3050% fewer
32 when a Salesforce Product2 connection is active.
Upload turn tools
51 → 2649% fewer
Upload turns keep preflight paths, not normal catalog mutation tools.
All flags on
72 → 4242% fewer
The widest surface still has snapshot fallbacks.

2.1 Tool-path metrics

Metric Baseline After Reviewer read
Tool definitions / round ~17.0k tokens ~11.6k tokens Prompt budget freed before model reasoning starts.
Tool calls / turn 2.06 1.69 Fewer tool round trips for the same jobs.
Round-1 reads / turn 0.71 0.26 Injected context is replacing first-round lookup calls.
Reads before first write 0.86 0.44 The agent gets to staging/writing sooner.
Latency p50 / p90 9.4s / 17.0s 9.3s / 16.6s Slightly better, not the main claimed win.
Spend per full run $10.19 $8.98 Lower suite cost from smaller tool schemas and fewer calls.
Attempt pass@1 93% 95% No functionality regression in the reported run.

2.2 Consolidation map

Area Old surface New surface Important semantics preserved or fixed
Approval rules 15 tools findApprovalRules, saveApprovalRules, proposeApprovalEdge, detectConflicts, evaluateTestContext Batch validates before write, patch updates start from locked stored rule, copyOf clones rules, missing create trigger can be inferred.
Approval groups 8 tools saveApprovalGroup, getApprovalGroupDetails, deleteApprovalGroup Omitted members keeps current assignees; replaceAllMembers means complete replacement, managers included.
Catalog staging Create/update/price/delete variants bulkStageCatalogProducts Eval expectations move to the batched ChangeSet tool name while executor compatibility remains for old persisted proposals.
Pricebooks Create/update split savePricebookForAgent, getPricebookForAgent, deletePricebookForAgent getPricebookForAgent stays because degraded snapshots cannot list a pricebook’s products.
Users and config state Name/id/list variants and separate draft reads findUsers, getDraftStatus Draft status, version history, and catalog drafts fail independently.
Salesforce Product2 Always-eligible catalog tools Product2 tools bind only when connection exists Normal orgs no longer pay schema tokens for tools they cannot use.

2.3 New context blocks

Turn starts
Runtime config includes org, user, conversation, and current upload keys.
Build snapshot
Adds approval groups, draft-aware approval rules, and composite price kinds.
Build files block
Lists this conversation’s uploads newest first, capped at 30, with exact R2 keys.
Bind tools
Registry drops tools made redundant by context and gates Product2 tools by connection.
Measure path
Telemetry records bound tool count, schema size, and per-round tool calls.
Org snapshot adds
  • Approval groups with member counts.
  • Approval rules, draft-aware, with live destination group names.
  • Degraded approval summaries when full rule text would exceed budget.
  • Composite price kinds in use, including truncated-catalog warnings.
Files context adds
  • Current turn attachments always listed.
  • Later turns get FILES IN THIS CONVERSATION.
  • Metadata only; no file contents injected.
  • readUploadedFile accepts exact key as well as filename.

2.4 Removed code and retained compatibility

Deletion shape

The PR deletes the replaced bound tools, including bulkWriteTools.ts, and removes a dead orchestrator guard. The PR description calls out −1,950 lines from replaced tools and net −665 lines in agent source.

Deleted: obsolete bound wrappers Kept: internal logic wrappers call Kept: snapshot fallback tools Kept: old ChangeSet executor names

3. How it works

3.1 Tool binding is now context-sensitive

Registry behavior

3.2 Approval rule saves are batch-first and patch-safe

Batch validation

saveApprovalRules plans the whole create/update/remove batch under the draft lock, rejects duplicate/conflicting IDs, loads validation data once, and writes nothing if any rule in the batch is invalid.

saveApprovalRules({
  create: [...],
  update: [...],
  remove: [...]
})
Patch semantics

Updates start from the locked stored rule. Omitted routes, trigger, and description remain exactly as saved, so stale agent reads do not restore old routing.

{ id: "edge-a", description: "Rename only" }
// keeps current trigger, routes, managedTab,
// and unknown future fields
Edge cases explicitly covered

3.3 Approval group saves no longer accidentally wipe members

Input shape Meaning after this PR Why reviewers should care
{ groupId, name } Settings-only patch; all members stay. Fixes the live-group member wipe risk.
{ members: [...] } Replace listed user members, preserve managers. Matches card-style submits that do not enumerate manager approvers.
{ members: [...], replaceAllMembers: true } Complete replacement; managers are included only if explicitly listed. Supports “make Jane the only approver.”
{ chainOrder } on create Allowed. Preserves chain placement creation behavior.
{ chainOrder } on update Refused. Avoids surprising reorder semantics on patch updates.

3.4 Eval harness now measures the route, not just the destination

Tool call log

Eval telemetry records per-round tool names from tool_round. Per-call status is not available yet, so failure-rate metrics report n/a instead of pretending zero failures.

Classifier

TOOL_KINDS classifies bound tools as read, check, write, UI, or plan. A spec builds the real registry with flags enabled and fails if any bound tool is unclassified.

Advisory expectations

Cases can assert path expectations such as notCalls, maxRound1Reads, or maxCallsOf. They report misses by default and only fail when promoted with gate: true.

ifAsked

A case may provide a scripted follow-up if the agent asks a clarifying question. The harness answers and grades the combined turn, so asking can be allowed without failing the scripted flow.

4. What it does not change

Explicit non-goals and preserved behavior

5. Tests, rollout signal, risks, and open questions

5.1 Test coverage added

New e2e step

E2E suite apps/server/e2e_tests/configAgent/ runs consolidated tools against a real database with throwaway organizations.

23/23reported passing
Unit coverage

New or expanded specs cover approval rule planning, approval group member merge logic, org snapshots, conversation files, draft status, registry gating, telemetry, tool path metrics, and report rendering.

Eval coverage

New regression cases pin custom-attribute deletion after approval and non-standard CSV headers ending with an actionable question/form/proposal instead of a dead end.

5.2 Eval status

Reported full run
57/61 cases pass 9 dataset skips 2,834/2,835 deterministic checks 77/78 judge checks No new regressions vs main

5.3 Risks and rollback

Primary risk: consolidated tools can accidentally change old edge-case semantics while preserving the common case. The PR mitigates this with targeted e2e tests and review fixes for patch behavior, member merge behavior, cloned approval rules, live group names, and independent draft/version reads.
Watch closely
  • Approval rule patch updates after concurrent edits.
  • Rules with invalid legacy conditions receiving description-only edits.
  • Manager preservation vs full replacement in approval groups.
  • Snapshot truncation paths for large catalogs and many approval rules.
  • File reads where duplicate filenames exist but keys differ.
Rollback shape
  • Revert the registry consolidation to re-bind old tools.
  • Keep old executor names available during rollback to avoid breaking persisted ChangeSets.
  • Use the new tool-path report to confirm bound tool counts and round-1 reads return to baseline if reverted.

5.4 Open questions and follow-ups

Not blockers unless reviewers disagree