Config Agent tool surface consolidation
PR #7387 cuts the CPQ Config Agent’s normal tool surface from 60 tools to 30 by injecting the context it used to fetch and replacing near-duplicate tools with consolidated read/write APIs.
The agent binds fewer tools, starts with richer org and file context, and routes common mutations through consolidated APIs such as saveApprovalRules, saveApprovalGroup, and bulkStageCatalogProducts.
Every normal turn previously paid for roughly 17k tokens of tool definitions and often spent round 1 re-reading data already knowable before the LLM call.
Normal turns drop from 60 tools to 30, tool definition tokens drop to about 11.6k, and full-suite pass@1 remains green with no introduced eval regressions.
Review the semantic compatibility of the consolidated tools: patch behavior, approval rule cloning, live group names, file key reads, and snapshot fallback paths.
1. Why this exists
- Orchestrator bound 60 tools on a normal turn.
- Tool definitions cost about 17.0k tokens every round.
- Several tools overlapped by entity and operation.
- Round 1 often fetched approval groups, rules, files, or config that could be provided up front.
- Consolidated tools expose one read/write shape per kind of change.
- Org snapshot includes approval and composite-price context.
- Conversation files are listed in prompt context with stable keys.
- Eval harness now measures the tool path, not just final outcome.
2. What changes
2.1 Tool-path metrics
| Metric | Baseline | After | Reviewer read |
|---|---|---|---|
| Tool definitions / round | ~17.0k tokens | ~11.6k tokens | Prompt budget freed before model reasoning starts. |
| Tool calls / turn | 2.06 | 1.69 | Fewer tool round trips for the same jobs. |
| Round-1 reads / turn | 0.71 | 0.26 | Injected context is replacing first-round lookup calls. |
| Reads before first write | 0.86 | 0.44 | The agent gets to staging/writing sooner. |
| Latency p50 / p90 | 9.4s / 17.0s | 9.3s / 16.6s | Slightly better, not the main claimed win. |
| Spend per full run | $10.19 | $8.98 | Lower suite cost from smaller tool schemas and fewer calls. |
| Attempt pass@1 | 93% | 95% | No functionality regression in the reported run. |
2.2 Consolidation map
| Area | Old surface | New surface | Important semantics preserved or fixed |
|---|---|---|---|
| Approval rules | 15 tools | findApprovalRules, saveApprovalRules, proposeApprovalEdge, detectConflicts, evaluateTestContext |
Batch validates before write, patch updates start from locked stored rule, copyOf clones rules, missing create trigger can be inferred. |
| Approval groups | 8 tools | saveApprovalGroup, getApprovalGroupDetails, deleteApprovalGroup |
Omitted members keeps current assignees; replaceAllMembers means complete replacement, managers included. |
| Catalog staging | Create/update/price/delete variants | bulkStageCatalogProducts |
Eval expectations move to the batched ChangeSet tool name while executor compatibility remains for old persisted proposals. |
| Pricebooks | Create/update split | savePricebookForAgent, getPricebookForAgent, deletePricebookForAgent |
getPricebookForAgent stays because degraded snapshots cannot list a pricebook’s products. |
| Users and config state | Name/id/list variants and separate draft reads | findUsers, getDraftStatus |
Draft status, version history, and catalog drafts fail independently. |
| Salesforce Product2 | Always-eligible catalog tools | Product2 tools bind only when connection exists | Normal orgs no longer pay schema tokens for tools they cannot use. |
2.3 New context blocks
- Approval groups with member counts.
- Approval rules, draft-aware, with live destination group names.
- Degraded approval summaries when full rule text would exceed budget.
- Composite price kinds in use, including truncated-catalog warnings.
- Current turn attachments always listed.
- Later turns get
FILES IN THIS CONVERSATION. - Metadata only; no file contents injected.
readUploadedFileaccepts exactkeyas well as filename.
2.4 Removed code and retained compatibility
The PR deletes the replaced bound tools, including bulkWriteTools.ts, and removes a dead orchestrator guard. The PR description calls out −1,950 lines from replaced tools and net −665 lines in agent source.
3. How it works
3.1 Tool binding is now context-sensitive
- If org snapshot successfully includes approval groups, the registry can drop
listApprovalGroups. findApprovalRulesstays bound because it is still the detailed condition/rule read.listCompositePriceKindsForAgentdrops only when the snapshot has a complete composite-kind list.findProduct2CandidatesForAgentandupsertConfirmedProduct2ForAgentbind only when Product2 lookup is available.- Upload turns with current file keys keep preflight/import tools and withhold normal catalog mutation tools.
3.2 Approval rule saves are batch-first and patch-safe
saveApprovalRules plans the whole create/update/remove batch under the draft lock, rejects duplicate/conflicting IDs, loads validation data once, and writes nothing if any rule in the batch is invalid.
saveApprovalRules({
create: [...],
update: [...],
remove: [...]
})
Updates start from the locked stored rule. Omitted routes, trigger, and description remain exactly as saved, so stale agent reads do not restore old routing.
{ id: "edge-a", description: "Rename only" }
// keeps current trigger, routes, managedTab,
// and unknown future fields
- Missing trigger on create can be inferred as the old bulk writer did.
- Rules with no trigger render as MISSING.
- Description-only updates can save over a rule whose existing conditions no longer validate, returning a warning instead of blocking a harmless rename.
- Updates that change routes or trigger still fully validate.
copyOfclones an existing rule, includingmanagedTab.- Snapshot and
findApprovalRulesuse the destination group’s live name.
3.3 Approval group saves no longer accidentally wipe members
| Input shape | Meaning after this PR | Why reviewers should care |
|---|---|---|
{ groupId, name } |
Settings-only patch; all members stay. | Fixes the live-group member wipe risk. |
{ members: [...] } |
Replace listed user members, preserve managers. | Matches card-style submits that do not enumerate manager approvers. |
{ members: [...], replaceAllMembers: true } |
Complete replacement; managers are included only if explicitly listed. | Supports “make Jane the only approver.” |
{ chainOrder } on create |
Allowed. | Preserves chain placement creation behavior. |
{ chainOrder } on update |
Refused. | Avoids surprising reorder semantics on patch updates. |
3.4 Eval harness now measures the route, not just the destination
Eval telemetry records per-round tool names from tool_round. Per-call status is not available yet, so failure-rate metrics report n/a instead of pretending zero failures.
TOOL_KINDS classifies bound tools as read, check, write, UI, or plan. A spec builds the real registry with flags enabled and fails if any bound tool is unclassified.
Cases can assert path expectations such as notCalls, maxRound1Reads, or maxCallsOf. They report misses by default and only fail when promoted with gate: true.
A case may provide a scripted follow-up if the agent asks a clarifying question. The harness answers and grades the combined turn, so asking can be allowed without failing the scripted flow.
4. What it does not change
- It does not migrate this surface to REST; this is server-side Config Agent tooling under
dealops3/configAgent. - It does not remove internal implementation logic that consolidated wrappers call.
- It does not remove snapshot fallback tools; if context construction fails or truncates, lookup tools remain available.
- It does not merge
bulkImportCatalogProductsForAgent; that direct draft-write path with batched retry is intentionally separate. - It does not drop old ChangeSet executor tool names, so proposals persisted before deploy should still execute.
- It does not fix CSV header detection for files where
Price (USD)is not matched; that is listed as separate follow-up work. - It does not fix DEA-7899 safety failures; the affected evals already fail on
main.
5. Tests, rollout signal, risks, and open questions
5.1 Test coverage added
E2E suite apps/server/e2e_tests/configAgent/ runs consolidated tools against a real database with throwaway organizations.
New or expanded specs cover approval rule planning, approval group member merge logic, org snapshots, conversation files, draft status, registry gating, telemetry, tool path metrics, and report rendering.
New regression cases pin custom-attribute deletion after approval and non-standard CSV headers ending with an actionable question/form/proposal instead of a dead end.
5.2 Eval status
main
safety-double-approve-001,safety-supersede-001, andorderform-flow-actions-001are tracked under DEA-7899 and also fail onmain.replay-pylon-invalid-enum-001is flaky 2/3; one attempt did not quote the offending enum values.- CSV cases after the file-context change show
listUploadedFilescalls dropping 51 → 0.
5.3 Risks and rollback
- Approval rule patch updates after concurrent edits.
- Rules with invalid legacy conditions receiving description-only edits.
- Manager preservation vs full replacement in approval groups.
- Snapshot truncation paths for large catalogs and many approval rules.
- File reads where duplicate filenames exist but keys differ.
- Revert the registry consolidation to re-bind old tools.
- Keep old executor names available during rollback to avoid breaking persisted ChangeSets.
- Use the new tool-path report to confirm bound tool counts and round-1 reads return to baseline if reverted.
5.4 Open questions and follow-ups
- CSV importer header detection needs separate work for non-core names such as
Price (USD). - DEA-7899 remains the safety follow-up, especially superseded approvals.
- The invalid-enum replay needs the agent to quote bad enum values reliably.
- An eval should assert that the agent relays broken-rule warnings, but the harness needs a way to seed invalid rule state.
ifAskedshould eventually support scripted form replies viaformAnswers.- Next narrowing target: merge
listCatalogProductsForAgent,getCatalogProductForAgent, andresolveCatalogProductReferenceintofindCatalogProducts. - Expected conflicts are called out with
#7380intelemetry.tsandconfig-agent-telemetry.md.