HR Handbook For Operators 20 min read Updated May 2026

Supervising AI Workers

You don't sit behind a human employee's shoulder watching every keystroke. You don't sit behind your AI employee's shoulder either. You set up the right alerts, read the digests, and look at the audit log when something feels off. Proportional oversight, dropping from 30 min/day in week 1 to 5 min/day by week 3.

This is the supervision rhythm. The trust ladder. The operator's playbook for managing 5 AI workers — or 50. Plus: how to supervise work you don't understand (legal, medical, technical), how multi-operator supervision works without dropped balls, what auditors want to see, and the seven supervision mistakes from the field that operators most often hit.

Quick Answer

Supervision drops from 30 min/day in week 1 to ~5 min/day by week 3. Three layers: (1) live watching via the agent's terminal/browser stream when you don't trust it yet, (2) daily digests in Slack / email summarizing yesterday's runs, (3) alerts on failure for steady state. If the trajectory isn't dropping, the agent isn't trustworthy yet — fix the CLAUDE.md, not the supervision. For 50-agent fleets, aggregate into one rolled-up dashboard rather than 50 separate digests.

Why supervise AI workers at all?

Three reasons supervision matters even after the shakedown:

  1. Drift detection. Agents drift. The CLAUDE.md you wrote in week 1 may not cover the edge cases that emerge in week 8. The agent that was perfect in March may be missing a new vendor portal's new button by June. Supervision catches drift before it becomes systemic.
  2. Quality assurance. The agent's output looks fine in summary but might be wrong in detail. Spot-checks catch quality issues that the agent's own self-reporting misses. The audit log is great evidence; spot-checks are great signal.
  3. Compliance and accountability. Regulators, auditors, customers, and partners want to know that a human is in the loop. "We supervise our AI workers daily, with these specific artifacts" is the answer that holds up under scrutiny. "We hired an AI and assumed it would work" doesn't.

The trap to avoid: supervision-as-anxiety (refreshing the dashboard every 20 minutes because you don't trust the agent). Anxiety supervision is exhausting and signals that the agent isn't ready for production. The right response to anxiety isn't more watching; it's tightening the CLAUDE.md and the guardrails until you can step back without dread.

The trust ladder: from week 1 to steady state

Supervision is proportional to trust, and trust is earned over time. The same is true of human employees — you read every email a new hire sends in week 1, but by month 3 you only read the ones they cc you on. The same trajectory works for AI:

PhaseTime / day per agentWhat you doWhat signals readiness for next phase
Week 1 — Shakedown 20-30 min Live-watch every run. Approve or correct. Update CLAUDE.md as gaps emerge. 5 consecutive days of clean runs; 2+ CLAUDE.md updates absorbed; you can predict what the agent will do
Week 2 — Calibration 5-10 min Read the daily digest. Spot-check 1-2 runs. Add to CLAUDE.md if you see drift. Digests reading routine; spot-checks reveal no surprises; alert-rate dropping
Week 3+ — Steady state 2-5 min Alerts on failure only. Read the weekly summary. Quarterly performance review. (this is the destination — no further phase)
Drift detected Bump back to week 1 Drop back to live-watching until you understand why behavior changed. Same as above — re-establish baseline

The key insight: the trajectory should be downward. If after three weeks you're still spending 30 min/day per agent, something's wrong — usually one of three things:

One supervision misfit we see often: operators who never escalate from week-1 watching, then complain "AI agents take too much time to manage." They're stuck in week 1 by choice, not by necessity. The trust ladder works when you let yourself climb it.

Layer 1: Live watching (week 1, and any time you intervene)

The dashboard exposes a real-time view of every agent. For coding agents (Claude Code, Codex CLI, Devin), you see the terminal — every command, every output, the conversation between agent and OS. For browser agents (Anthropic Computer Use, Operator), you see Chrome live with the cursor moving. For desktop agents, you see the XFCE desktop. It's the equivalent of pulling up a chair next to a human's monitor on day one.

What you're checking for in week 1

The shadow-an-employee analogy holds

Think about how you'd shadow a human contractor on day one. You're not micromanaging — you're calibrating. After 3-5 runs, you have a sense of where the gaps are. Patch them in the CLAUDE.md and step back. The same shape works for AI workers, just compressed from 2 weeks to 2 hours.

What live-watching looks like by role

RoleWhat you'd watchWhat "right" looks like
AI AP clerk (Nina) The Slack channel as she processes invoices; the NetSuite UI to verify each posting; her summary message at end Each invoice extracted correctly, posted to staging (not GL), Slack post is concise, escalations flagged with reason
AI paralegal (Amy) The terminal as she searches Westlaw; the draft memo as it forms; her citation list Citations verify against database, memo follows the firm's template, escalations cite specific clause types
AI research analyst (Robin) The browser as she visits competitor pages; the Notion doc as she drafts the brief Each source visited, accurate quotes, structured comparison, source URLs cited
AI coding agent (Ada) The terminal as Claude Code runs; the PR as it forms; CI status Diffs scoped to the task, tests pass, PR description follows the template
AI customer-support tier-1 (Sam) The ticket queue as Sam triages; her draft replies before they send; the escalation log KB-search lands the right article, draft is on-brand, ambiguous tickets escalated
AI practice manager (Patty) The payer portals as she logs in; the EHR as she updates eligibility; her flag log No PHI leaks beyond BAA boundary, eligibility codes correct, flagged exceptions clear

Watching multiple runs simultaneously

For operators with 3-5 agents, the dashboard's "split view" shows multiple live streams at once. You're not staring at any one for long; you're scanning for anything that catches your eye. Most operators we work with set up:

This is the cockpit view. 10 minutes a day in this layout covers what would take 45 minutes of switching between dashboards.

Layer 2: The daily digest (week 2 onward)

By week 2, you don't have time to live-watch. The agent should be posting a daily summary that you skim in 30 seconds. Set this up in the CLAUDE.md (the job description) — something like:

"At 5pm every weekday, post a daily summary to #nina-summary. Format: invoices processed (count + total $), invoices flagged for review (count + reason), vendors not responding (list), and an estimate of tomorrow's expected workload. Include a link to the run log for any flagged item."

Why digests work

  1. They surface anomalies you'd otherwise miss in a sea of green runs. "We processed 50 invoices today; one was flagged for $42K vendor amount" — that's the line you read first.
  2. The agent's own framing of what happened is often more readable than the raw audit log. Plain-English summaries beat 200 audit-log entries every time.
  3. Aggregate trends become visible. Volume going up, error rate creeping up, vendors going dark — the digest exposes these before they become incidents.
  4. Your team can read it, not just you. Distributed supervision: the AP-clerk digest is in #ap-summary; the controller, the bookkeeper, and you all see it.
  5. The digest is itself an audit artifact. If a regulator asks "what did your AI worker do in Q3?", the daily digest archive is the answer.

Digest formats by role

Concrete templates that work in production:

Slack — AP clerk daily digest example
📊 *Nina — AP Daily Summary, Tue Mar 11*

✅ Processed: 47 invoices ($142,830 total)
⚠️ Flagged: 3 (links: #1247, #1251, #1268)
   • #1247 — vendor unrecognized; first time processing this vendor
   • #1251 — amount above $10K threshold; needs Marcus approval
   • #1268 — PO not found in NetSuite; portal showed PO# 88471
🔁 Vendors not responding: Acme (3 days), Globex (1 day)
📅 Tomorrow expected: ~50 invoices (Wed is high-volume day)
🔗 Full run log: link.example.com/run/abc123
Slack — Paralegal daily digest example
📊 *Amy — Paralegal Daily, Tue Mar 11*

📄 NDAs triaged: 8 (5 standard, 2 mutual, 1 non-standard flagged)
   • Non-standard: BlackPearl Capital (term clause unusual; partner review)
🔍 Case research completed: 3 memos delivered
   • Smith v Jones (NY commercial) — 5 cases cited, all verified
   • In re Acme bankruptcy — 12 cases cited, all verified
   • State v Roe (CA appellate) — 7 cases cited, all verified
⚖️ Filings: 2 (motion to dismiss filed; reply brief filed)
📅 Tomorrow: discovery review for 3 cases
🔗 Full run log: link.example.com/run/def456
Slack — Customer support tier-1 daily digest example
📊 *Sam — Support Daily, Tue Mar 11*

🎫 Tickets handled: 142 total
  • Auto-resolved: 89 (62%)
  • Escalated to tier-2: 34 (24%)
  • Awaiting customer reply: 19 (14%)
📈 Volume vs avg: +12% (Tuesday is high)
⚠️ Anomalies:
  • 7 tickets about login issue → likely auth incident, FYI
  • Customer NPS sentiment slightly down (3 negative replies)
📅 Backlog at midnight: 28 tickets
🔗 Full run log: link.example.com/run/ghi789

The digest format is in the CLAUDE.md. Tighten it over time as you learn what's signal and what's noise.

Digest cadence

Role typeCadenceWhy
Event-driven (AP clerk, paralegal, support)Daily at end-of-businessOperator reads at coffee the next morning
Bursty (research, code review)End-of-batchReviews tied to actual work, not artificial date boundaries
Long-running (always-on monitoring)Weekly with daily heartbeatDaily activity is low-info; weekly trends are signal
Cron-triggered batch (overnight competitive intel)Per-batchOne digest after each scheduled run
Mission-critical (financial agents)Daily + immediate alerts on thresholdEnd-of-day for review; immediate for anything material

Layer 3: Alert on failure (steady state)

By week 3, you want the agent to ping you only when it can't proceed or when it failed. Set up Slack alerts for:

Crucially: no alert for a successful run. Successful runs go in the digest. Alerts are reserved for things that need your attention. If you're getting too many alerts, your alert criteria are wrong — tighten them. If you're getting too few alerts and finding issues yourself, the criteria are too loose.

Alert tiering: not every alert is a page

TierCriteriaChannelResponse time
P0 (page)Customer-affecting; cost runaway; security eventPagerDuty / phoneImmediate
P1 (urgent)Hard failure during business hours; compliance alertSlack DM with @hereWithin 1 hour business
P2 (review)Quality flag; anomaly; escalationSlack channelWithin 4 hours business
P3 (FYI)Threshold crossed but not urgent (budget at 60%)Slack channelDaily digest

Most operator alerts should be P2 or P3. P0/P1 should be rare; if you're getting more than 2-3 per month per agent, the agent has bigger issues than supervision can fix.

The audit log: when you need to dig

Every run leaves an audit trail. The dashboard's audit-log view shows:

You'll only look at the audit log when something feels off — a customer complains the AP clerk paid an invoice twice; a partner asks why the paralegal's brief cited the wrong case; an auditor wants to see what the agent did with patient data in Q3. The audit log is your forensics.

Retention

Keep audit logs for as long as your compliance regime requires:

Sandbox Platform's audit log retention is configurable per workspace. Default is 90 days; for regulated workloads, set it to 7 years and let the platform's lifecycle archival handle it.

For routine investigation, the dashboard's audit view supports:

For compliance-grade access, the audit log API lets you query programmatically; the platform also supports streaming the audit feed to your SIEM (Splunk, Datadog, Sumo).

Real-time watching vs async review

You'll do both, but the ratios shift over time:

Real-time watching

  • Use when: debugging, training a new agent, investigating a complaint
  • Cost: high (your full attention)
  • Yields: precise feedback, immediate intervention
  • Frequency: Week 1 default; rare by week 3
  • Tool: dashboard live stream

Async review

  • Use when: daily/weekly check-ins, performance review, audit
  • Cost: low (5-10 min)
  • Yields: trends, anomaly detection, confidence
  • Frequency: Weekly default; daily during shakedown
  • Tool: Slack digest, audit log, dashboard summaries

The healthy ratio for a steady-state agent: 5% real-time, 95% async. For a new agent or one in trouble: invert it — 50/50 or higher real-time during shakedown.

Supervising long-running agents (the always-on pattern)

Some agents — research agents, on-call agents, monitoring agents — never stop. Their workspace persists across days, accumulating state. Supervision for these is different:

Read more in The Always-On Agent and Persistent Workspaces.

Supervising event-driven agents

Event-driven agents (the ones that wake up on webhooks, cron, calendar entries) have a particular failure mode: silent dead-letter. An event arrived, no agent picked it up, nobody noticed. Specific safeguards:

When you have 50 agents instead of 5

Supervising 50 agents is not 10× supervising 5. With the right setup, it's 1.5×. Strategies:

Aggregate dashboards instead of per-agent attention

Don't read 50 daily digests. Read one rolled-up dashboard view:

That's your daily 5-minute glance. The agents that need attention bubble up; the rest you trust until they don't.

One unified digest channel

Don't have 50 Slack channels — have one rolling digest channel where every agent posts. Skim it daily; the volume is manageable because each agent posts one summary, not a dozen detail messages. Use threading: each agent's daily summary is one parent message; details go in the thread.

Tier supervision by criticality

TierExamplesSupervision posture
Mission-criticalCustomer-facing support; financial-posting agents; regulatory filingsTight alerts, daily review, monthly perf review
High-impact internalAP clerk; QA tester; code reviewerStandard alerts, daily digest, quarterly review
Low-impact internalResearch / drafting / weekly digestsLoose alerts, weekly check, quarterly review
Experimental / R&DPrototypes; new role onboardingTreat like week 1 — live-watch every run

Most fleets we work with end up with 4-7 mission-critical agents getting tight supervision and the other 30-40 in the standard tier. Trying to give all 50 the mission-critical treatment is what burns out operators.

Distributed supervision: not one operator, many

For larger fleets, supervision is a team activity. Patterns that work:

When to intervene (not just supervise)

Sometimes you have to step in. The signals:

Intervening means: pause the agent → read the audit log → fix the CLAUDE.md or the tool access → resume. It is NOT: take over the task yourself. Your job is to make the agent better, not to do the agent's job.

The intervention anti-pattern

Operators new to AI workforce management sometimes intervene reflexively. The agent does something slightly different from how the operator would; operator pauses and corrects; rinse, repeat. Two months in, the operator is doing 60% of the work themselves and complaining the agent isn't pulling its weight.

The fix: intervene only when the result is wrong, the cost is high, the customer is affected, or a hard rule was broken. Stylistic differences aren't grounds for intervention. The agent's job is to produce correct outputs; how it gets there is its business.

Seven supervision mistakes from the field

Anonymized but real. The most common patterns operators new to AI workforce management hit:

1. Anxiety supervision

Operator refreshes the dashboard every 20 minutes, even when the agent's been running clean for weeks. Burns the operator out; signals to the team that the AI workforce isn't trustworthy. Fix: pick a metric (5 consecutive clean weeks) and commit to dropping to digest-review when you hit it. The agent has earned the trust; the operator has to grant it.

2. No primary supervisor

Three operators all assume someone else is reading the digest. Nobody is. Three weeks in, an issue surfaces that should have been caught in week 1. Fix: assign one primary per agent; rotate quarterly. The other operators can route work to the agent (in Slack), but the supervisor catches drift.

3. Alert fatigue

Operator set up too many alerts in week 1 — every escalation, every flagged item, every cost threshold. After a month they're auto-archiving the alert channel. Real issues now go unnoticed. Fix: tier alerts (P0/P1/P2/P3); only P0/P1 hit Slack; P2/P3 in the digest. Re-tune monthly.

4. Supervision-as-substance

Operator who can't verify the legal substance of the AI paralegal's work decides to "review" it anyway, scanning briefs they don't really understand. They miss substance issues; they miss process issues too because they're playing the wrong game. Fix: split substance review (senior in-domain human, monthly spot-check) from process supervision (operator, daily digest).

5. Treating disagreement as drift

Agent does the task correctly but in a way the operator wouldn't have. Operator pauses, corrects, intervenes. Three weeks of this and the operator is doing more work than before. Fix: intervene only on wrong results, not on stylistic differences. If the way matters, encode it in CLAUDE.md once.

6. No audit-log discipline

Two years in, a regulator asks "what did this agent do with our customer data in Q3 last year?" Operator can't produce the audit trail because the workspace's audit retention was set to 90 days. Fix: set audit retention to your compliance regime's required length on day one. Storage is cheap; missing audit is expensive.

7. Confusing supervising with intervening

Operator's "supervision time" is actually intervention time — they're stopping the agent, taking over the task, finishing it themselves. Cost is 3× baseline because everything is touched twice. Fix: separate the two activities. Supervision is reading. Intervention is doing. Track them separately.

Cross-shift handoffs

For agents that operate across time zones or 24/7, supervision needs handoff discipline. Pattern that works:

What an auditor wants to see

SOC2, HIPAA, SOX, GDPR, FedRAMP — every audit regime has a slightly different set of artifacts they want to see. The common ones for AI worker supervision:

  1. Audit log of every action. Sandbox Platform captures this; ensure retention matches your regime.
  2. Supervision-rhythm proof. Alerts received and resolved; digests reviewed (timestamps); reviews completed. The platform's supervision-history view exports this as a CSV.
  3. CLAUDE.md history. When did the rules change? Why? Use git for the CLAUDE.md if you can; the platform supports versioning natively.
  4. Escalation log. When did the agent stop and ask? Who answered, what was the resolution? This proves human-in-the-loop.
  5. Performance-review records. Quarterly reviews with metrics. Read the Performance Reviews guide for the framework.
  6. Incident records. When something went wrong, what happened, what was the root cause, what changed to prevent recurrence.
  7. Access controls. Who has dashboard access; who can modify CLAUDE.md; who can pause/terminate agents. Workspace permissions in the platform.
  8. Encryption and residency posture. If applicable to your regime — proof that data stays where it's supposed to and is encrypted at rest.

Auditors generally don't care about the AI specifics; they care about the same things they always care about — controls, audit trails, change management, access controls. Frame your supervision evidence in those terms.

Supervision and customer trust

If your AI workers touch customer data or customer-facing surfaces, supervision matters externally too. Three patterns:

Disclosure

Tell customers an AI worker is in the loop. Most companies in 2026 default to disclosure ("AI assists; humans review and approve") because the alternative — getting caught not disclosing — is much worse. The platform supports configurable customer-facing labels (e.g. "Drafted by Nina (AI), reviewed by Marcus").

Right-to-human

Some industries (some banking, some healthcare, some legal) require that customers can request a human at any time. Bake this into the agent's CLAUDE.md as a hard rule: "if the customer says 'I want to speak with a human' in any phrasing, immediately escalate to #human-support."

Recourse and remediation

If the agent makes a mistake that affects a customer, you need a clear remediation path: how to investigate (audit log), how to fix (rollback or correction), how to compensate (refund, credit, apology), how to prevent (CLAUDE.md update). Document this once; reuse it.

Supervision-time benchmarks by role

Anonymized benchmarks from operators we work with at steady state (post week 3):

RoleTime / day per agent (steady state)Time / week per agent
AP clerk2-3 min15-20 min + 30 min weekly review
Paralegal5-7 min30-45 min + 60 min weekly review (legal substance)
Research analyst3-4 min20-30 min + 30 min weekly review
Code reviewer5-8 min30-45 min (engineers review the agent's reviews)
Customer support tier-14-6 min30-40 min + 30 min weekly NPS review
Practice manager (HIPAA)6-10 min45-60 min + monthly compliance review
Always-on monitoring1-2 min (heartbeat)5-10 min weekly state review
QA tester3-4 min20-30 min

Aggregate for a 5-agent fleet: ~30 min/day. For a 10-agent fleet with mixed roles: ~45-60 min/day across whichever operators are doing supervision. For a 50-agent fleet with proper tier-2 distribution: ~2-3 hours/day across the whole supervision team.

Frequently asked questions

How much time should I spend supervising an AI employee?

Week 1: 20-30 min/day per agent (live watching). Week 2: 5-10 min/day (digest review). Week 3+: 2-5 min/day (alerts only). The trajectory is dropping; if it isn't, the agent isn't trustworthy enough yet — fix the CLAUDE.md, not the supervision.

What's the difference between supervising and intervening?

Supervising is passive: alerts, digests, audit logs. Intervening is active: stopping a run, correcting a step, taking over the task. Most of your time should be supervising; intervening is the exception. Once you find yourself intervening more than once a week per agent, the job description (CLAUDE.md) needs updating, not your time.

How do I know if an agent has gone rogue or is just slow?

Three signals: (1) cost-per-task suddenly higher than baseline, (2) escalation rate dropping to zero (it should ask for help sometimes), (3) result quality dipping in spot-checks. If any two are true, intervene. If only one is anomalous, watch closer for a few days; could be normal variance.

Why does the agent post in Slack instead of just emailing me?

Threading. Slack lets you reply to a specific run with feedback, and the agent reads the thread. "Nina — that vendor portal needs MFA, I'll set that up" goes in the thread; the agent picks it up and acts. Email is one-way; Slack is conversational. For high-volume work, Slack scales much better — and the threading captures the back-and-forth as audit-log-with-narrative.

Can I supervise from my phone?

Yes — Slack and the dashboard's mobile view. Most operators we talk to do their morning supervision pass on the train, not at a desk. The whole point of "5 min/day" supervision is that it should fit in the gaps between meetings.

What about agents working overnight while I'm asleep?

That's the whole point — they work while you sleep, you read the digest at coffee. The night-shift agent posts results at 7am local; you skim it before standup. If something needed your attention, the alert woke you. Otherwise, results are waiting.

How do I supervise an agent doing work I don't understand?

This is the hard one. If your AI paralegal cites case law you can't verify, you can't supervise the substance — you can only supervise the process (did it cite cases? did it format correctly? did it escalate edge cases?). For substance, you'll want a senior human in that domain to do periodic spot-checks. AI workers don't eliminate the need for domain expertise; they redirect it from doing-the-work to reviewing-the-work.

What if I disagree with how the agent did something?

Update the CLAUDE.md, not the model. The job description is where you encode your preferences. "Always cite the Bluebook 21st edition." "Use our company's invoice template, not the vendor's." "Sign emails 'The Sandbox Team', not 'AI Assistant'." Write it down once; the agent reads it every run.

How does supervision work when multiple humans share an AI worker?

Pick a primary supervisor — one human who owns the CLAUDE.md and reviews the daily digest. Other operators can still route tasks to the agent (in Slack), but the supervisor catches drift. Without a primary, supervision falls between the cracks; everybody assumes someone else is watching.

What does an auditor want to see during a supervision review?

Five artifacts: (1) the audit log of every action the agent took, (2) the supervision-rhythm proof (alerts received, digests reviewed), (3) the CLAUDE.md history (when did rules change, why), (4) the escalation log (when did the agent stop and ask, who answered), (5) the performance-review records. The platform's audit log captures most of this automatically; the CLAUDE.md history is your responsibility (use git or the platform's version-control).

How quickly should I respond to an alert?

Depends on the alert tier. Hard failures (agent crashed, downstream system rejected): review within an hour during business hours, next morning otherwise. Cost-spike alerts: real-time (something is wrong). Quality flags: review within a day. Anomaly alerts (escalation rate dropped to zero): review within 48 hours. Calibrate so alerts feel meaningful, not noisy.

How do I supervise 50 agents without a full-time team?

Aggregate. Don't try to read 50 daily digests; read one rolled-up dashboard view that surfaces only the agents that need attention. Per-agent quarterly performance reviews (not weekly); aggregate alerts from all agents into one Slack channel; tier supervision by criticality (mission-critical agents get tighter alerts, internal-only agents run looser).

What's the supervision rhythm for an always-on agent?

Different shape. Live-watching doesn't apply — there's no "task start" to watch. Use heartbeats (agent posts I'm-alive every N hours), drift checks (snapshot the agent's state weekly and review), and restart-on-schedule (monthly even if nothing seems wrong, to clear accumulated cruft). Read more in The Always-On Agent.

Can I see what the agent typed character-by-character?

Yes — for terminal agents (Claude Code, Codex CLI, Devin), the dashboard shows a live terminal stream. For browser agents, you see Chrome's screen in real time via the access_url. For desktop agents, the XFCE desktop. This is the "pull up a chair next to a new hire on day one" equivalent. Use it during the shakedown; don't make a habit of it past week 1.

What if an agent does the right thing but in a way I wouldn't have?

Don't intervene. The agent learning a different-but-correct path is fine; you're paying for results, not method. Intervene only if the result is wrong, the cost is high, the customer is affected, or the method violates a hard rule (compliance, safety, brand voice). Stylistic differences are not a supervision issue.

What's next

Manage your AI workforce

The dashboard's supervision view shows every agent at a glance. 5 minutes a day for 50 workers.

Open Dashboard