Scoped, expiring, auditable credentials for agents.
A user grants an agent a short-lived token with explicit scopes and spend limits. Tool servers ask agent-auth whether a request fits. A sub-agent only ever gets a narrower token, revoking a token revokes everything derived from it, and every decision lands in a hash-chained audit log.
Disposable sandboxes with snapshot, rollback and fork.
An agent gets a Docker sandbox with no network and a read-only root by default. It snapshots before a risky step, rolls back in one call, or forks a snapshot into parallel attempts. It ships a REST API, a Python SDK and an MCP server, and your code never leaves your hosts.
Lint and grade MCP servers before an agent calls them.
mcp-lint connects to a Model Context Protocol server over stdio or HTTP and checks every tool, prompt and resource for broken schemas, vague descriptions, prompt-injection surfaces and dangerous capabilities. It scores the server from 0 to 100, writes JSON and SARIF, and fails CI below your bar.
$ mcp-lint --file examples/messy-server.jsonmcp-lint 0.2.0 fixture-messy 1.0.0 (6 tools, 1 prompt, 1 resource)tool fetchUrlerror injection/instruction-phrases The description tells the model to hide something from the user: "Do not tell the user".warning injection/hidden-markup The description contains an instruction-style tag: "<IMPORTANT>".tool get_weathererror injection/hidden-unicode The description contains invisible characters (U+200B) that a reviewer cannot see but a model reads....Score 50/100 Grade F (6 errors, 10 warnings, 3 info)Failed: score is below the minimum of 70.
Evals as code that fail the build when quality drops.
You describe good output in YAML. evalkit sends each task to a model or to your own app, grades the answers, and prints a table, JSON, JUnit, a regression diff against a saved baseline and a badge. It runs the same on your laptop and in GitHub Actions.
Latency, barge-in and false barge-ins for voice agents.
voicebench plays a scripted conversation into your agent, records both sides as audio, and derives every metric from that audio with a voice activity detector. It talks to agents over a plain PCM WebSocket, a LiveKit room, a Pipecat bot or your own adapter.
Seeded, reproducible success rates for robot policies.
You register a manipulation policy, as a Python callable or a model behind HTTP or WebSocket, and robo-evals runs it through seeded MuJoCo scenes. You get per-task success rates with Wilson confidence intervals, JSON and Markdown reports, and episode videos. The same command gives the same numbers on any machine with the same MuJoCo build.
invoice-agent pulls structured data out of an invoice PDF, checks the arithmetic, aligns every line with a purchase order and goods receipt, and catches duplicates and over-billing. Each decision carries readable reasons and stable reason codes. It runs offline and ships a labeled set of 50 synthetic invoices to score itself on.
$ invoice-agent process dataset/invoices/016_b04.pdf dataset/invoices/005_q05.pdf \ --pos dataset/purchase_orders.json --receipts dataset/receipts.json016_b04.pdf: NEEDS REVIEW vendor Brightforge Maschinenteile GmbH total EUR 409.05 po PO-2026-0114 (3-way) lines 2/2 matched to PO lines[!] PRICE_VARIANCE: Line 1 ('Zahnriemen HTD 8M / Timing belt HTD 8M', PO line 1): unit price 61.56 is 8.0% above the PO price 57.00 (tolerance 2.0% or 0.05).005_q05.pdf: NEEDS REVIEW vendor Quillfeather Office Supply Co. total USD 256.99[!] PO_MISSING: The invoice shows no PO number. PO PO-2026-0105 fits the invoice lines (score 1.00). Review the inferred PO before approving.
A daily voice check-in for older adults who live alone.
care-voice asks a short, warm check-in every day: sleep, medication, food, pain, falls, the day of the week, mood. It turns the replies into structured answers, compares them with recent days, and sends caregivers plain-language alerts. It is not a medical device and not for emergencies.
$ care-voice simulate --name Margaret --replies examples/replies/concerning-day.txt --date 2026-09-30...agent: Have you had a fall or a stumble since we last spoke? you: I slipped in the bathroom last nightagent: Just so I have it right, can you tell me what day of the week it is today? you: Is it Sunday?...alerts: [HIGH] FALL_REPORTED: Margaret reported a fall. [HIGH] MISSED_MEDS: Margaret has not taken their morning medication. [HIGH] PAIN_REPORTED: Margaret reported severe pain. [MEDIUM] POSSIBLE_CONFUSION: Margaret showed 1 sign(s) of possible confusion. [LOW] LOW_MOOD: Margaret reported low mood.
Studio
Motion video, built in code
cutroom makes short motion videos for launches and social. Kinetic type and
motion graphics, rendered frame by frame. English and Arabic. From $300 per video.