California's No Robo Bosses Act: Human-in-the-Loop AI Rules
California just banned AI-only firing and discipline decisions. Here is what the No Robo Bosses Act means and how to build human-in-the-loop AI.
We've spent the last 11 months shipping voice agent deployments for coaches, consultants, fintech, real estate, and a handful of edge cases. Ninety-six in production. Here's what we've learned about what actually works in 2026.
1. The model isn't the bottleneck anymore
GPT-4o-realtime, Claude 3.5 Sonnet voice, and the open-source equivalents are good enough for 92% of production scenarios. Telephony latency, audio processing pipelines, and prompt routing are now the failure modes not LLM quality.
If your agent feels janky, audit your audio path before you audit your prompts. Eight times out of ten, that's where the friction lives.
"The agents that work feel like infrastructure. The agents that fail feel like party tricks."
2. Voice ≠ chatbot with audio
Every team that tries to port their chatbot prompt to voice fails the same way: too verbose, too formal, too explainer-y. Voice is improv. You need shorter turns, callback handles, and graceful interruption.
3. The handoff is the product
The best voice agent in the world is useless if the post-call sync is broken. Notes go to CRM. CRM triggers sequence. Sequence books follow-up. Calendar invites human. That is the system. The voice piece is one component.
If you want to see a live example, our AI calling system is running in production for loan servicing and collections you can see the real numbers on the case studies page.
On September 30, California Governor Gavin Newsom signed SB 947, the No Robo Bosses Act. It bars employers from relying solely on automated systems to fire or discipline workers. It is the first law of its kind in the United States, and it reverses a veto Newsom issued on an earlier version of the bill in 2025. If you run AI agents anywhere near people decisions, human-in-the-loop AI just moved from best practice to legal requirement.
Even if you are not in California, read on. Laws like this spread. The design pattern they require is also simply how reliable AI systems get built.
What the No Robo Bosses Act actually requires
The law is narrower than the headlines suggest, and that matters for how you respond. It does not ban AI in HR. It targets automated decision systems that drive termination and discipline outcomes. The core requirements reported so far:
- No AI-only decisions. If an employer relies primarily on AI output to fire or discipline someone, a human reviewer must corroborate it using other information, such as manager evaluations, peer reviews, and personnel files.
- Written notice. The affected employee must be told AI was primarily used, given a description of the employee data the system looked at, and given a human contact who can explain the decision.
- Delayed effect. The law is reported to become operative on July 1, 2027, so there is a window to prepare.
Notice what is not in there: a ban on scoring, flagging, or summarizing. Software can still surface performance patterns. A person just has to own the decision, and the record has to show they actually looked.
Why this is bigger than HR
The legal trigger is employment, but the underlying principle is general: consequential decisions about people need a traceable human checkpoint. You can already see the same logic showing up in other places:
- Debt collection rules that limit what automated outreach can say and when.
- The EU AI Act, which classifies employment and credit decisions as high risk.
- Insurance, lending, and tenant screening rules that require adverse action explanations.
If your business uses AI to score leads, prioritize collections accounts, or screen applicants, you are one regulation away from the same question: who reviewed this, and what did they see? Teams that answer it early spend a weekend on it. Teams that answer it late spend a quarter on it.
The three-part pattern for human-in-the-loop AI
Good human oversight is not a rubber stamp button. A reviewer who approves 400 items an hour is not reviewing anything, and a regulator will see that in the logs. Build it as three parts.
1. Tier decisions by consequence
Not every action needs a human. Sorting an inbox does not. Terminating a contract does. Define three tiers: AI acts alone (reversible, low stakes), AI drafts and a human approves (medium stakes), and AI only recommends while a human decides with independent evidence (high stakes, anything touching a person's income, housing, credit, or job).
2. Force corroboration, not just approval
The statute language about corroborating evidence is the useful design hint. Do not show the reviewer only the AI conclusion. Show the underlying records the AI used plus at least one source the AI did not use. If the reviewer cannot disagree with the model using other evidence, the review is cosmetic.
3. Log everything an auditor would ask for
Store the model output, the input data, the reviewer identity, the time spent, the evidence viewed, and the final decision with a reason. When the notice requirement applies, you need to describe what data the system used. That is trivial if you logged it at decision time and nearly impossible to reconstruct afterward.
How to build the checkpoint in n8n or any workflow tool
This is easier than it sounds. In n8n, the pattern is a workflow that runs the AI step, writes the result and its inputs to a database, then pauses on a wait or approval node and notifies a reviewer in Slack or email with a link to a review screen. The workflow resumes only when a named person submits a decision with a reason. Everything is written to an audit table. We covered the mechanics in our post on human-in-the-loop agents in n8n.
Three details separate a real checkpoint from theater:
- Timeouts that escalate, not auto-approve. If nobody responds, the item goes to a manager, never through.
- Reviewer capacity limits. If the queue exceeds what a person can sensibly review, the system slows intake instead of lowering the bar.
- Disagreement tracking. Measure how often reviewers override the AI. A rate near zero means rubber stamping. A very high rate means your model is not ready.
The same structure applies in voice and collections. At Nexica, we have handled $48.9M in accounts with AI calling systems that are TCPA compliant by design, and the pattern is always the same: automation does the volume work, humans own the exceptions and the consequential calls. See how this works in our AI agents and workflow automation builds.
What to do this quarter
You have until mid 2027 for California, but the work is cheap now and expensive later. A practical checklist:
- List every place AI output influences a decision about a person: hiring, scoring, discipline, collections, pricing, eligibility.
- Assign each one a consequence tier.
- For tier three, add a mandatory human review step with corroborating evidence and a written decision reason.
- Start logging inputs and outputs today so your history is complete when rules take effect.
- Draft the notice language and name the human contact now, so it is a template and not a scramble.
The companies that win with AI over the next two years will not be the ones that remove humans from every loop. They will be the ones that put humans in exactly the right loops and can prove it.
If you want this built for your business, book a 20-minute call with Nexica AI. We build production-grade AI systems in 14 days.