Skip to main content
Moveworks agents run inside the Moveworks cloud (Agent Studio + Reasoning Engine), so there is no in-process callback to hook the way LangChain or CrewAI integrations do - Moveworks Python Script Actions have outbound network access disabled at the infrastructure level. Instead, this integration pulls activity out through the Moveworks Data API (the read-only OData export at https://api.moveworks.ai/export/v1) and replays it into AgentX: Every imported trace is stamped with the moveworks framework label, so Moveworks traffic gets its own series in Monitor’s Platforms chart and the Live Traces framework filter. No extra dependency is needed - the importer uses only the SDK’s core requirements.

Prerequisites

  • Data API credentials, minted by a Moveworks superadmin (Moveworks Setup → Credentials). API key and OAuth2 client-credentials are both offered there; the importer takes the key and sends it as Authorization: Bearer <key> (override the header/scheme if your deployment differs).
  • An AgentX project API key (self-host: Settings → API access).

CLI: scheduled sync

The SDK installs an agentx-moveworks command built for cron. It keeps an incremental cursor (~/.agentx/moveworks_cursor.json by default), so each run picks up where the last one ended:
Useful flags: --dry-run prints the mapped payloads without ingesting, --agent-name attributes every trace to one agent instead of per-domain, --timestamp-field overrides the record timestamp field the OData $filter uses (default created_time), and --no-cursor ignores the cursor file. On any ingest failure the cursor is not advanced, so a re-run retries the window. Three flags control evaluation of the imported traffic:
  • Pattern/built-in checks (PII, empty response, tool failure, active patterns) and trace-scoped online evaluators run on every imported trace automatically, same as any live traffic - no flag needed (--monitor is kept for compatibility).
  • --judge-sessions judges every imported session with each enabled session-scoped evaluator (Session Baseline Judge included) after the sync. Needed for backfills: the engine’s automatic session sweep only looks at the last 24 hours of activity, so older imported conversations would otherwise never get a session verdict. The requests carry ifStale=true, so a session the sweep (or a previous run) already scored is skipped, never judged twice. Each judgment is a real LLM call - budget accordingly on large backfills.
  • --evaluate-against <dataset_or_config_id> grades each new imported interaction’s recorded input/output against that grading config’s criteria - see Evaluating Moveworks agents below.

Python API

Evaluating Moveworks agents

Moveworks agents run inside the Moveworks cloud and can’t be invoked from outside (no inbound converse API; Script Actions have outbound network disabled). That constrains how each kind of evaluation applies - but both kinds work: Online evaluation - fully supported. Every synced interaction rides the normal ingest pipeline: trace-scoped online evaluators sample and score it, patterns raise signals, and --judge-sessions gets whole conversations judged (goal progression, consistency, resolution). One caveat to hold honestly: the Data API has a ~24h freshness SLA, so this is continuous retroactive scoring - yesterday’s production judged today - not at-the-moment scoring. Signals, triage, calibration, and judge tuning all apply unchanged. Offline evaluation - grade recorded behavior instead of re-running the agent. A classic dataset run calls your agent per case; with no way to invoke a Moveworks agent, the offline paths work from what the agent already did:
  • --evaluate-against <id> grades each imported interaction against a grading config during the sync (one judge call per new interaction - the engine’s span dedupe means re-syncing a window never re-bills). Each verdict is recorded as a real one-result evaluation run, so ratings and justifications show up under Evaluate → Runs, and the report prints the average:
  • The same thing per trace from code: client.tracer.evaluate_trace(trace_id, dataset_id).
  • Datasets from production: any synced trace or conversation converts to a golden dataset case (multi-turn included) - the regression suite for grading future synced traffic, or the benchmark for evaluating a candidate replacement.
  • Model portability: replay a synced interaction’s input against other models (client.monitor.run_model_portability(trace_id, [...])) for a cost/quality comparison.
  • If your deployment exposes any inbound HTTP entry point, the http_endpoint adapter closes the loop with real dataset runs.

Behavior and limits

  • Idempotent: every trace carries a deterministic span_id (mw:<interaction_id>) and the engine skips spans it has already ingested, so re-syncing a window never duplicates.
  • Historical timestamps are honored: traces land in AgentX dated when the conversation happened, not when it was imported, so the Overview windows, cost chart buckets, and session timelines attribute correctly.
  • Near-real-time, not live: the Data API follows a ~24-hour freshness SLA and retains 30 days - an hourly or daily sync is the honest cadence.
  • No token counts: Moveworks does not export usage, so these traces contribute nothing to the LLM cost chart.
  • No fabricated failures: plugin calls carry served/used flags, not success/failure - the importer records those flags verbatim and never marks a tool call failed, so Moveworks traffic cannot generate false tool-failure signals.
  • Deployment-tolerant field mapping: Data API field names vary between deployments and API versions; the importer reads each field from a list of known candidates and skips only records with no parseable timestamp (counted in the report as skipped_no_time).