From d4207d478a3a74583960e7b1d1d84fbc91199862 Mon Sep 17 00:00:00 2001 From: jordanc-relevanceai Date: Mon, 7 Sep 2026 10:47:42 +1000 Subject: [PATCH] docs(operations): move Evals and Version history into a new Operations section Evals covers Agents and Workforces; version history covers Agents, Tools, and Workforces. Both lived under Agents. They now sit in an Operations section between Invent and Agents in the Build with Relevance AI tab. - Evaluate Performance keeps its dropdown - Version history is a single page, so it stays a single page - 6 redirects cover the moved URLs, plus 3 existing /evals shortcuts repointed - Inbound links updated repo-wide Co-Authored-By: Claude Opus 5 (1M context) --- build/agents/create-an-agent.mdx | 2 +- build/invent/invent.mdx | 4 +- build/{agents => operations}/evals/checks.mdx | 4 +- .../evals/introduction.mdx | 2 +- .../{agents => operations}/evals/monitor.mdx | 2 +- .../evals/running-evaluations.mdx | 4 +- .../evals/test-sets.mdx | 6 +- .../version-history.mdx | 0 build/tools/create-a-tool.mdx | 2 +- build/workforces/create-a-workforce.mdx | 2 +- changelog.mdx | 2 +- docs.json | 57 ++++++++++++++----- get-started/pricing.mdx | 2 +- guides/customer-success/account-health.mdx | 2 +- guides/customer-success/getting-started.mdx | 2 +- guides/customer-success/qbr-prep.mdx | 2 +- guides/customer-success/renewal-expansion.mdx | 2 +- guides/customer-support/getting-started.mdx | 2 +- guides/customer-support/kb-generation.mdx | 2 +- guides/customer-support/response-drafting.mdx | 2 +- guides/customer-support/ticket-triage.mdx | 2 +- guides/marketing/campaign-analytics.mdx | 2 +- guides/marketing/content-repurposing.mdx | 2 +- guides/marketing/getting-started.mdx | 2 +- guides/marketing/lifecycle-campaigns.mdx | 4 +- guides/revops/data-dedup.mdx | 2 +- guides/revops/getting-started.mdx | 2 +- guides/revops/lead-routing.mdx | 2 +- guides/revops/pipeline-hygiene.mdx | 2 +- guides/sales/competitive-intelligence.mdx | 2 +- guides/sales/crm-data-enrichment.mdx | 2 +- guides/sales/getting-started.mdx | 2 +- guides/sales/lead-scoring.mdx | 2 +- guides/sales/meeting-briefs.mdx | 2 +- guides/sales/personalized-outbound.mdx | 2 +- guides/sales/prospect-research-use-case.mdx | 2 +- guides/solution-engineering/demo-prep.mdx | 2 +- .../discovery-summaries.mdx | 2 +- .../solution-engineering/getting-started.mdx | 2 +- guides/solution-engineering/rfp-responses.mdx | 2 +- 40 files changed, 87 insertions(+), 58 deletions(-) rename build/{agents => operations}/evals/checks.mdx (97%) rename build/{agents => operations}/evals/introduction.mdx (97%) rename build/{agents => operations}/evals/monitor.mdx (91%) rename build/{agents => operations}/evals/running-evaluations.mdx (92%) rename build/{agents => operations}/evals/test-sets.mdx (95%) rename build/{agents/build-your-agent => operations}/version-history.mdx (100%) diff --git a/build/agents/create-an-agent.mdx b/build/agents/create-an-agent.mdx index 4a3f6677..52939395 100644 --- a/build/agents/create-an-agent.mdx +++ b/build/agents/create-an-agent.mdx @@ -88,4 +88,4 @@ Once you've created an Agent, explore the rest of the build guides to refine it: - [Triggers](/build/agents/build-your-agent/triggers) — run your Agent automatically on a schedule, from a webhook, or from an integration - [Alerts](/build/agents/build-your-agent/alerts) — get notified when your Agent needs attention, and define when it should loop in a human - [Memory](/build/agents/build-your-agent/memory) — configure how your Agent retains context across conversations -- [Version history](/build/agents/build-your-agent/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions +- [Version history](/build/operations/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions diff --git a/build/invent/invent.mdx b/build/invent/invent.mdx index e93a6e45..97cd6e10 100644 --- a/build/invent/invent.mdx +++ b/build/invent/invent.mdx @@ -220,7 +220,7 @@ Invent is available to organization admins by default. An organization admin can Ask Invent to roll an asset back to an earlier version — for example, "Restore the previous version of this tool". Invent copies that version into the current draft, leaving the live published version unchanged until you publish. This restores an earlier version of an existing asset; it doesn't recover one that's been deleted. - Invent always asks you to confirm before restoring, and appends `(restored)` to the version name so the restored state is easy to identify later. For a full guide to version history — including how to restore and rename versions from the builder UI, and how retention works — see [Version history](/build/agents/build-your-agent/version-history). + Invent always asks you to confirm before restoring, and appends `(restored)` to the version name so the restored state is easy to identify later. For a full guide to version history — including how to restore and rename versions from the builder UI, and how retention works — see [Version history](/build/operations/version-history). @@ -230,7 +230,7 @@ Invent is available to organization admins by default. An organization admin can Invent can build an initial Evals suite, review results, investigate failed Checks, and diagnose performance alarms for Agents and Workforces. - + Learn how to create scenarios and reusable Checks, run evaluations, and monitor live performance. diff --git a/build/agents/evals/checks.mdx b/build/operations/evals/checks.mdx similarity index 97% rename from build/agents/evals/checks.mdx rename to build/operations/evals/checks.mdx index 6a0ee458..18a112a9 100644 --- a/build/agents/evals/checks.mdx +++ b/build/operations/evals/checks.mdx @@ -10,7 +10,7 @@ A Check is one pass-or-fail criterion. Every Agent and Workforce has its own lib - **To a Monitor dashboard** — it runs on a sampled portion of live tasks. - **To a one-off evaluation** of already-completed tasks selected from the task list. -This page is the reference for what each type scores. To attach one while building a Test, see [creating Tests](/build/agents/evals/test-sets). +This page is the reference for what each type scores. To attach one while building a Test, see [creating Tests](/build/operations/evals/test-sets). --- @@ -99,7 +99,7 @@ The Checks tab lists every Check on the Agent or Workforce, grouped by type. Fil --- -Next: Learn how to [run Evals](/build/agents/evals/running-evaluations) against your Test sets. +Next: Learn how to [run Evals](/build/operations/evals/running-evaluations) against your Test sets. ## Frequently asked questions (FAQs) diff --git a/build/agents/evals/introduction.mdx b/build/operations/evals/introduction.mdx similarity index 97% rename from build/agents/evals/introduction.mdx rename to build/operations/evals/introduction.mdx index abb029d6..6d83065b 100644 --- a/build/agents/evals/introduction.mdx +++ b/build/operations/evals/introduction.mdx @@ -89,7 +89,7 @@ Clicking a **Credits** or **Actions** value opens a breakdown of where that Test --- -Next: Learn how to [create Tests](/build/agents/evals/test-sets) — the simulated conversations your Agent gets evaluated against. +Next: Learn how to [create Tests](/build/operations/evals/test-sets) — the simulated conversations your Agent gets evaluated against. ## Frequently asked questions (FAQs) diff --git a/build/agents/evals/monitor.mdx b/build/operations/evals/monitor.mdx similarity index 91% rename from build/agents/evals/monitor.mdx rename to build/operations/evals/monitor.mdx index 6568f9eb..8840925a 100644 --- a/build/agents/evals/monitor.mdx +++ b/build/operations/evals/monitor.mdx @@ -4,7 +4,7 @@ sidebarTitle: 'Monitoring Evals' description: 'Score the real conversations your Agent or Workforce is having against Checks, on dashboards you configure per sample rate' --- -The **Monitor** section continuously scores live tasks against [Checks](/build/agents/evals/checks). Unlike the **Test** section, which runs simulated conversations, Monitor evaluates the real conversations your Agent or Workforce is having. +The **Monitor** section continuously scores live tasks against [Checks](/build/operations/evals/checks). Unlike the **Test** section, which runs simulated conversations, Monitor evaluates the real conversations your Agent or Workforce is having. Monitor is organized into **dashboards** — you can create more than one (for example, one focused on tone, another on tool-use accuracy) and configure each independently. diff --git a/build/agents/evals/running-evaluations.mdx b/build/operations/evals/running-evaluations.mdx similarity index 92% rename from build/agents/evals/running-evaluations.mdx rename to build/operations/evals/running-evaluations.mdx index 29d274ea..5ca6a84d 100644 --- a/build/agents/evals/running-evaluations.mdx +++ b/build/operations/evals/running-evaluations.mdx @@ -4,7 +4,7 @@ sidebarTitle: 'Running Evals' description: 'Run Test sets and individual Tests, read the scores that come back, and require evals to pass before publishing' --- -Once you've created [Tests](/build/agents/evals/test-sets) and defined [Checks](/build/agents/evals/checks), you're ready to run evaluations and see how your Agent or Workforce performs. You can run them by hand, or have a publish run them for you. +Once you've created [Tests](/build/operations/evals/test-sets) and defined [Checks](/build/operations/evals/checks), you're ready to run evaluations and see how your Agent or Workforce performs. You can run them by hand, or have a publish run them for you. --- @@ -94,7 +94,7 @@ Once configured, click **Save**. When you next publish your Agent, the selected --- -Next: Learn about [monitoring Evals](/build/agents/evals/monitor) — scoring the real conversations your Agent or Workforce is having. +Next: Learn about [monitoring Evals](/build/operations/evals/monitor) — scoring the real conversations your Agent or Workforce is having. ## Frequently asked questions (FAQs) diff --git a/build/agents/evals/test-sets.mdx b/build/operations/evals/test-sets.mdx similarity index 95% rename from build/agents/evals/test-sets.mdx rename to build/operations/evals/test-sets.mdx index b1b29eae..ae4ec82f 100644 --- a/build/agents/evals/test-sets.mdx +++ b/build/operations/evals/test-sets.mdx @@ -4,7 +4,7 @@ sidebarTitle: 'Creating Tests' description: 'Create Tests that simulate real user conversations with your Agent or Workforce, and group them into Test sets' --- -A **Test** simulates one conversation with your Agent or Workforce. You describe a user in the **Scenario** field — *"You are a long-time customer who was charged twice for the same order"* — and Relevance plays that user, generates their messages, and hands back the transcript. The [Checks](/build/agents/evals/checks) you attach score it. +A **Test** simulates one conversation with your Agent or Workforce. You describe a user in the **Scenario** field — *"You are a long-time customer who was charged twice for the same order"* — and Relevance plays that user, generates their messages, and hands back the transcript. The [Checks](/build/operations/evals/checks) you attach score it. Tests live in a **Test set**, which is what a run points at — run the whole set, or a single Test inside it. @@ -47,7 +47,7 @@ Tests live in a **Test set**, which is what a run points at — run the whole se - See [Check types](/build/agents/evals/checks) for what each type scores and the fields it takes. + See [Check types](/build/operations/evals/checks) for what each type scores and the fields it takes. 8. Set **When the test hits an approval** to control how the Test resolves Tool approvals and escalations mid-conversation: @@ -112,7 +112,7 @@ Tests can be reorganized across Test sets as your testing strategy evolves. Each --- -Next: Read the [Check types](/build/agents/evals/checks) reference — what each type scores and how the Checks tab works. +Next: Read the [Check types](/build/operations/evals/checks) reference — what each type scores and how the Checks tab works. ## Frequently asked questions (FAQs) diff --git a/build/agents/build-your-agent/version-history.mdx b/build/operations/version-history.mdx similarity index 100% rename from build/agents/build-your-agent/version-history.mdx rename to build/operations/version-history.mdx diff --git a/build/tools/create-a-tool.mdx b/build/tools/create-a-tool.mdx index 386b6981..d015a417 100644 --- a/build/tools/create-a-tool.mdx +++ b/build/tools/create-a-tool.mdx @@ -99,4 +99,4 @@ Once you've created a Tool, explore the rest of the build guides to refine it: - [Inputs](/build/tools/customize-tool/inputs) — configure what data your Tool receives - [Steps](/build/tools/customize-tool/steps) — chain actions together into a workflow - [Outputs](/build/tools/customize-tool/outputs) — define what your Tool returns -- [Version history](/build/agents/build-your-agent/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions +- [Version history](/build/operations/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions diff --git a/build/workforces/create-a-workforce.mdx b/build/workforces/create-a-workforce.mdx index 562f14af..3429699a 100644 --- a/build/workforces/create-a-workforce.mdx +++ b/build/workforces/create-a-workforce.mdx @@ -42,4 +42,4 @@ Once you've created a Workforce, explore the rest of the build guides to refine - [Add Tools](/build/workforces/build-an-ai-workforce/add-tools) — connect Tools directly within your workflow - [Add conditions](/build/workforces/build-an-ai-workforce/add-conditions) — create routing rules for smarter workflows - [Edge settings](/build/workforces/build-an-ai-workforce/edge-settings) — configure how connections between nodes behave -- [Version history](/build/agents/build-your-agent/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions +- [Version history](/build/operations/version-history) — every save creates a snapshot you can restore, so you can experiment without losing working versions diff --git a/changelog.mdx b/changelog.mdx index 30e0edef..109347d0 100644 --- a/changelog.mdx +++ b/changelog.mdx @@ -42,7 +42,7 @@ Eval runs now show exactly where their cost goes, itemized per check, and cover Each completed Eval run renders a cost breakdown panel that itemizes credit and action consumption per component. Checks account for 1 action per run across every check type, including LLM as Judge and Tool Usage, so the panel's totals reconcile to the exact source of every credit and action. Workforce runs that fan out to sub-agents or call tools record and evaluate each of those interactions, the same as individual Agent runs. -To see the breakdown, open any completed Eval run; the panel appears alongside the results. See the [Evals documentation](/build/agents/evals/introduction) for full details. +To see the breakdown, open any completed Eval run; the panel appears alongside the results. See the [Evals documentation](/build/operations/evals/introduction) for full details. diff --git a/docs.json b/docs.json index ddf85285..284af591 100644 --- a/docs.json +++ b/docs.json @@ -110,6 +110,22 @@ "build/invent/invent" ] }, + { + "group": "Operations", + "pages": [ + { + "group": "Evaluate Performance (Evals)", + "pages": [ + "build/operations/evals/introduction", + "build/operations/evals/test-sets", + "build/operations/evals/checks", + "build/operations/evals/running-evaluations", + "build/operations/evals/monitor" + ] + }, + "build/operations/version-history" + ] + }, { "group": "Agents", "pages": [ @@ -123,7 +139,6 @@ "build/agents/build-your-agent/alerts", "build/agents/build-your-agent/memory", "build/agents/build-your-agent/variables", - "build/agents/build-your-agent/version-history", { "group": "Trigger Types", "pages": [ @@ -167,16 +182,6 @@ } ] }, - { - "group": "Evaluating Performance", - "pages": [ - "build/agents/evals/introduction", - "build/agents/evals/test-sets", - "build/agents/evals/checks", - "build/agents/evals/running-evaluations", - "build/agents/evals/monitor" - ] - }, { "group": "Running Tasks", "pages": [ @@ -783,6 +788,30 @@ } }, "redirects": [ + { + "source": "/build/agents/evals/introduction", + "destination": "/build/operations/evals/introduction" + }, + { + "source": "/build/agents/evals/test-sets", + "destination": "/build/operations/evals/test-sets" + }, + { + "source": "/build/agents/evals/checks", + "destination": "/build/operations/evals/checks" + }, + { + "source": "/build/agents/evals/running-evaluations", + "destination": "/build/operations/evals/running-evaluations" + }, + { + "source": "/build/agents/evals/monitor", + "destination": "/build/operations/evals/monitor" + }, + { + "source": "/build/agents/build-your-agent/version-history", + "destination": "/build/operations/version-history" + }, { "source": "/enterprise/org-project-controls-governance", "destination": "/enterprise/manage-projects-and-users-api" @@ -1145,15 +1174,15 @@ }, { "source": "/evals", - "destination": "/build/agents/evals/introduction" + "destination": "/build/operations/evals/introduction" }, { "source": "/build/agents/evals", - "destination": "/build/agents/evals/introduction" + "destination": "/build/operations/evals/introduction" }, { "source": "/build/agents/build-your-agent/evals", - "destination": "/build/agents/evals/introduction" + "destination": "/build/operations/evals/introduction" } ] } diff --git a/get-started/pricing.mdx b/get-started/pricing.mdx index 16aa5ee4..c6622766 100644 --- a/get-started/pricing.mdx +++ b/get-started/pricing.mdx @@ -161,7 +161,7 @@ Start delegating work to agents today. Grow as your AI Workforce expands. | **Bring your own LLM** | ✗ | ✓ | ✓ | ✓ | | **A/B Testing** | ✗ | ✓ | ✓ | ✓ | | **[Analytics Dashboard](/enterprise/analytics)** | ✗ | ✗ | ✓ | ✓ | -| **[Agent Evaluations](/build/agents/evals/introduction)** | ✗ | ✗ | ✗ | ✓ | +| **[Agent Evaluations](/build/operations/evals/introduction)** | ✗ | ✗ | ✗ | ✓ | | **Work Hour Controls** | ✗ | ✗ | ✗ | ✓ | ### Security & compliance diff --git a/guides/customer-success/account-health.mdx b/guides/customer-success/account-health.mdx index 51a57200..091f2e04 100644 --- a/guides/customer-success/account-health.mdx +++ b/guides/customer-success/account-health.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Pipe churn outcomes back into the Agent's [evals](/build/agents/evals/introduction) so scoring tracks what predicted cancellation. + Pipe churn outcomes back into the Agent's [evals](/build/operations/evals/introduction) so scoring tracks what predicted cancellation. diff --git a/guides/customer-success/getting-started.mdx b/guides/customer-success/getting-started.mdx index ced90a17..2636e55e 100644 --- a/guides/customer-success/getting-started.mdx +++ b/guides/customer-success/getting-started.mdx @@ -105,7 +105,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Set up evals so health scoring or QBR-prep regressions get caught before they reach the CSM. - + Define test cases in the UI. Run them on every change. diff --git a/guides/customer-success/qbr-prep.mdx b/guides/customer-success/qbr-prep.mdx index 3bbb4952..b182c9b7 100644 --- a/guides/customer-success/qbr-prep.mdx +++ b/guides/customer-success/qbr-prep.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which QBR sections landed into the Agent's [evals](/build/agents/evals/introduction) so the format tracks what's used. + Feed back which QBR sections landed into the Agent's [evals](/build/operations/evals/introduction) so the format tracks what's used. diff --git a/guides/customer-success/renewal-expansion.mdx b/guides/customer-success/renewal-expansion.mdx index 54b11892..3a7cb8a2 100644 --- a/guides/customer-success/renewal-expansion.mdx +++ b/guides/customer-success/renewal-expansion.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed renewal outcomes back into the Agent's [evals](/build/agents/evals/introduction) so it leans on the plays that saved accounts. + Feed renewal outcomes back into the Agent's [evals](/build/operations/evals/introduction) so it leans on the plays that saved accounts. diff --git a/guides/customer-support/getting-started.mdx b/guides/customer-support/getting-started.mdx index edd13dc0..4ac52510 100644 --- a/guides/customer-support/getting-started.mdx +++ b/guides/customer-support/getting-started.mdx @@ -105,7 +105,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Set up evals so regressions in tone or accuracy get caught before the Agent replies to a customer. - + Define test cases in the UI. Run them on every change. diff --git a/guides/customer-support/kb-generation.mdx b/guides/customer-support/kb-generation.mdx index 1806f2c0..a72735a7 100644 --- a/guides/customer-support/kb-generation.mdx +++ b/guides/customer-support/kb-generation.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed read and deflection data back into the Agent's [evals](/build/agents/evals/introduction) so it drafts the formats that help. + Feed read and deflection data back into the Agent's [evals](/build/operations/evals/introduction) so it drafts the formats that help. diff --git a/guides/customer-support/response-drafting.mdx b/guides/customer-support/response-drafting.mdx index 4e0babe2..df51e530 100644 --- a/guides/customer-support/response-drafting.mdx +++ b/guides/customer-support/response-drafting.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which drafts shipped untouched, plus CSAT, into the Agent's [evals](/build/agents/evals/introduction) so it matches what your team sends. + Feed back which drafts shipped untouched, plus CSAT, into the Agent's [evals](/build/operations/evals/introduction) so it matches what your team sends. diff --git a/guides/customer-support/ticket-triage.mdx b/guides/customer-support/ticket-triage.mdx index c85563cf..b787adbc 100644 --- a/guides/customer-support/ticket-triage.mdx +++ b/guides/customer-support/ticket-triage.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed resolved-ticket outcomes back into the Agent's [evals](/build/agents/evals/introduction) so its triage tracks what actually held up. + Feed resolved-ticket outcomes back into the Agent's [evals](/build/operations/evals/introduction) so its triage tracks what actually held up. diff --git a/guides/marketing/campaign-analytics.mdx b/guides/marketing/campaign-analytics.mdx index 24a2781a..ba0634e6 100644 --- a/guides/marketing/campaign-analytics.mdx +++ b/guides/marketing/campaign-analytics.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Weight the Agent's [evals](/build/agents/evals/introduction) toward the metrics that actually predicted pipeline. + Weight the Agent's [evals](/build/operations/evals/introduction) toward the metrics that actually predicted pipeline. diff --git a/guides/marketing/content-repurposing.mdx b/guides/marketing/content-repurposing.mdx index 4a930faa..fe009648 100644 --- a/guides/marketing/content-repurposing.mdx +++ b/guides/marketing/content-repurposing.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed channel engagement back into the Agent's [evals](/build/agents/evals/introduction) so it leans on what performed. + Feed channel engagement back into the Agent's [evals](/build/operations/evals/introduction) so it leans on what performed. diff --git a/guides/marketing/getting-started.mdx b/guides/marketing/getting-started.mdx index 10fb51e8..22cb54bb 100644 --- a/guides/marketing/getting-started.mdx +++ b/guides/marketing/getting-started.mdx @@ -105,7 +105,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Set up evals so brand-voice regressions get caught before they ship to your list. - + Define test cases in the UI. Run them on every change. diff --git a/guides/marketing/lifecycle-campaigns.mdx b/guides/marketing/lifecycle-campaigns.mdx index e8e6370c..f240e9cd 100644 --- a/guides/marketing/lifecycle-campaigns.mdx +++ b/guides/marketing/lifecycle-campaigns.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed opens, clicks, and conversion back into the Agent's [evals](/build/agents/evals/introduction) so drafts track what converts. + Feed opens, clicks, and conversion back into the Agent's [evals](/build/operations/evals/introduction) so drafts track what converts. @@ -87,7 +87,7 @@ Once it's running, deepen it in three moves: The HubSpot trigger fires on the wrong field and the sequence drafts the wrong campaign. Gate triggers tightly and review what actually causes them to fire in week one. - A prompt tweak changes tone everywhere. Use [evals](/build/agents/evals/introduction) on a few canonical drafts to catch drift before it ships to subscribers. + A prompt tweak changes tone everywhere. Use [evals](/build/operations/evals/introduction) on a few canonical drafts to catch drift before it ships to subscribers. Tempting to skip marketing approval and queue directly. Don't — one bad subject line in production costs trust. Gate L3 with a Slack approval until you've watched it run for a quarter. diff --git a/guides/revops/data-dedup.mdx b/guides/revops/data-dedup.mdx index dbf1815d..13afde5b 100644 --- a/guides/revops/data-dedup.mdx +++ b/guides/revops/data-dedup.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed merge undos and disputes back into the Agent's [evals](/build/agents/evals/introduction) so its confidence thresholds track what's reliable. + Feed merge undos and disputes back into the Agent's [evals](/build/operations/evals/introduction) so its confidence thresholds track what's reliable. diff --git a/guides/revops/getting-started.mdx b/guides/revops/getting-started.mdx index e61df374..9e711f3f 100644 --- a/guides/revops/getting-started.mdx +++ b/guides/revops/getting-started.mdx @@ -105,7 +105,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Set up evals so routing or enrichment regressions get caught before they propagate across thousands of records. - + Define test cases in the UI. Run them on every change. diff --git a/guides/revops/lead-routing.mdx b/guides/revops/lead-routing.mdx index d3104f60..f3d37d0e 100644 --- a/guides/revops/lead-routing.mdx +++ b/guides/revops/lead-routing.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which routing held up into the Agent's [evals](/build/agents/evals/introduction) so edge-case handling tracks real outcomes. + Feed back which routing held up into the Agent's [evals](/build/operations/evals/introduction) so edge-case handling tracks real outcomes. diff --git a/guides/revops/pipeline-hygiene.mdx b/guides/revops/pipeline-hygiene.mdx index 32141a42..b32d4410 100644 --- a/guides/revops/pipeline-hygiene.mdx +++ b/guides/revops/pipeline-hygiene.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which at-risk signals predicted slippage into the Agent's [evals](/build/agents/evals/introduction) so flagging sharpens. + Feed back which at-risk signals predicted slippage into the Agent's [evals](/build/operations/evals/introduction) so flagging sharpens. diff --git a/guides/sales/competitive-intelligence.mdx b/guides/sales/competitive-intelligence.mdx index 855c4fce..b2cd1d61 100644 --- a/guides/sales/competitive-intelligence.mdx +++ b/guides/sales/competitive-intelligence.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed win/loss notes back into the Agent's [evals](/build/agents/evals/introduction) so it tracks the signals that matter. + Feed win/loss notes back into the Agent's [evals](/build/operations/evals/introduction) so it tracks the signals that matter. diff --git a/guides/sales/crm-data-enrichment.mdx b/guides/sales/crm-data-enrichment.mdx index 2c719f54..54aabcc0 100644 --- a/guides/sales/crm-data-enrichment.mdx +++ b/guides/sales/crm-data-enrichment.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Tune it through [evals](/build/agents/evals/introduction) so it maintains the fields your team actually uses. + Tune it through [evals](/build/operations/evals/introduction) so it maintains the fields your team actually uses. diff --git a/guides/sales/getting-started.mdx b/guides/sales/getting-started.mdx index 825fcc80..2f9842e5 100644 --- a/guides/sales/getting-started.mdx +++ b/guides/sales/getting-started.mdx @@ -104,7 +104,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Once you're shaping prompts and tools, set up evals so you catch regressions before customers do. - + Define test cases in the UI. Run them on every change. diff --git a/guides/sales/lead-scoring.mdx b/guides/sales/lead-scoring.mdx index 1901257c..6fc91b7e 100644 --- a/guides/sales/lead-scoring.mdx +++ b/guides/sales/lead-scoring.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed closed-won and closed-lost outcomes back into the Agent's [evals](/build/agents/evals/introduction) so "qualified" tracks what closes. + Feed closed-won and closed-lost outcomes back into the Agent's [evals](/build/operations/evals/introduction) so "qualified" tracks what closes. diff --git a/guides/sales/meeting-briefs.mdx b/guides/sales/meeting-briefs.mdx index 30cea4aa..19908834 100644 --- a/guides/sales/meeting-briefs.mdx +++ b/guides/sales/meeting-briefs.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which sections reps actually use into the Agent's [evals](/build/agents/evals/introduction) so briefs sharpen over time. + Feed back which sections reps actually use into the Agent's [evals](/build/operations/evals/introduction) so briefs sharpen over time. diff --git a/guides/sales/personalized-outbound.mdx b/guides/sales/personalized-outbound.mdx index 6e11709e..c45c6d7d 100644 --- a/guides/sales/personalized-outbound.mdx +++ b/guides/sales/personalized-outbound.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed reply and meeting rates back into the Agent's [evals](/build/agents/evals/introduction) so drafts work from what converts. + Feed reply and meeting rates back into the Agent's [evals](/build/operations/evals/introduction) so drafts work from what converts. diff --git a/guides/sales/prospect-research-use-case.mdx b/guides/sales/prospect-research-use-case.mdx index 9a3cde11..1ed65065 100644 --- a/guides/sales/prospect-research-use-case.mdx +++ b/guides/sales/prospect-research-use-case.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed won-deal signals back into the Agent's [evals](/build/agents/evals/introduction) so it tracks what closes. + Feed won-deal signals back into the Agent's [evals](/build/operations/evals/introduction) so it tracks what closes. diff --git a/guides/solution-engineering/demo-prep.mdx b/guides/solution-engineering/demo-prep.mdx index 22e4f850..c3217704 100644 --- a/guides/solution-engineering/demo-prep.mdx +++ b/guides/solution-engineering/demo-prep.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which brief sections SEs use into the Agent's [evals](/build/agents/evals/introduction) so briefs sharpen over time. + Feed back which brief sections SEs use into the Agent's [evals](/build/operations/evals/introduction) so briefs sharpen over time. diff --git a/guides/solution-engineering/discovery-summaries.mdx b/guides/solution-engineering/discovery-summaries.mdx index ca45b43b..cef9875c 100644 --- a/guides/solution-engineering/discovery-summaries.mdx +++ b/guides/solution-engineering/discovery-summaries.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which sections SEs reference into the Agent's [evals](/build/agents/evals/introduction) so summaries track what matters. + Feed back which sections SEs reference into the Agent's [evals](/build/operations/evals/introduction) so summaries track what matters. diff --git a/guides/solution-engineering/getting-started.mdx b/guides/solution-engineering/getting-started.mdx index 9f7edf5b..3c4ab2c0 100644 --- a/guides/solution-engineering/getting-started.mdx +++ b/guides/solution-engineering/getting-started.mdx @@ -105,7 +105,7 @@ How hands-on do you want to be? Pick a tab — the rest of the page is written f Set up evals so a wrong security answer or hallucinated architecture detail gets caught before it lands in front of a prospect. - + Define test cases in the UI. Run them on every change. diff --git a/guides/solution-engineering/rfp-responses.mdx b/guides/solution-engineering/rfp-responses.mdx index 1895b2c2..ab2c5ba3 100644 --- a/guides/solution-engineering/rfp-responses.mdx +++ b/guides/solution-engineering/rfp-responses.mdx @@ -73,7 +73,7 @@ Once it's running, deepen it in three moves: Wrap it in a [workflow](/build/workforces/create-a-workforce) that fires on a [trigger](/build/agents/build-your-agent/triggers). - Feed back which answers customers accepted into the Agent's [evals](/build/agents/evals/introduction) so it promotes the strong ones. + Feed back which answers customers accepted into the Agent's [evals](/build/operations/evals/introduction) so it promotes the strong ones.