Your AI wrote the code.
Assay tells you what's actually true.
One production app: 319 claims verified. 25 bugs. 8 of them were the spec's fault, not the code's.
Assay extracts every claim your app makes, verifies each against the code, then triages every failure. Real defect: fix it. Spec drift: leave the code alone.
A blind auto-fixer ships a security regression here. Assay tells your agent to stand down.
npx tryassay assess .thennpx tryassay fix .Runs with your Anthropic API key, roughly $1–10 in tokens by repo size. Reports publish to tryassay.ai by default; --no-publish keeps everything local.fix applies the real defects. Never the spec-drift ones.
We also point it at the big stuff.
OpenAI. Kubernetes. Shopify. Cloudflare. Vulnerabilities found and reported in all of them.
RCE via vm.Script sandbox escape
Buffer.from('').constructor.constructor('return process')()SSRF chain across kubeadm discovery endpoints
// RetrieveValidatedConfigInfo — httpsURL passed with no host validation client.Get(httpsURL) // user-controlled, full response consumed
Path traversal to root code execution in GCI mounter
path, _ := filepath.Split(os.Args[0]) // attacker controls argv[0]
rootfsPath := filepath.Join(path, rootfs)
// then: exec.Command("chroot", rootfsPath, "/bin/mount", ...)SSRF + CORS + Open Redirect chain
const remote = requestHeaders.get("X-CF-Remote");
await fetch(switchRemote(url, remote), { ... }) // no assertValidURL() callOAuth tokens written world-readable
writeFileSync(path.join(configPath), TOML.stringify(config), {
encoding: "utf-8", // missing mode option — defaults to 0o644 via umask
});Predictable CSP nonce via Math.random fallback
// Falls back to Math.random() when crypto.getRandomValues() throws return new Uint8Array(16).map(() => (Math.random() * 255) | 0);
Weak OAuth state parameter
const randomString = Math.random().toString(36).substring(2) // Same file correctly uses crypto.getRandomValues() elsewhere
CSRF wildcard origin bypass
// Guard intended to block *.com — but *.com has length 2, bypasses check if (patternParts.length === 1 && patternParts[0] === '**') return false
Not just big targets. Real developers. Real repos.
Drop a repo. We'll scan it free. Join the community →
Human-built code scores 91. Here's what AI platforms score.
4 platforms verified. 18 bugs found. 0 passed.
We publish the misses.
A blind holdout of real CVEs, re-scored from scratch every week. No cherry-picking.
of 164 blind holdout CVEs caught at the exact vulnerable code. Strictest definition.
points of blind recall gained last teaching cycle, same holdout.
real CVEs scored from the GitHub Advisory Database, each scanned pre-fix.
How it works
A neurosymbolic verification loop. Neural systems extract claims. Symbolic oracles verify them. The architecture that makes AI output trustworthy.
An LLM reads your code and identifies every implicit claim it makes. “This function handles null input.” “This API returns sorted results.” No regex or AST parser can do this — it requires semantic understanding. This is the neural layer.
A deterministic oracle tests each claim. For code, that means actual test execution in an isolated subprocess. The oracle doesn’t hallucinate. It runs the code and reports what happened. This is the symbolic layer.
Every failed claim is classified: real defect, spec drift, or uncertain. Real defects get a fix prompt your coding agent can apply, or let `assay fix` apply them for you. Spec drift gets flagged so nobody “fixes” correct code.
Claim extraction + dual-direction checking · U.S. Patent App. #63/980,048
Start free.
Try Assay on your code. No credit card required.
For solo developers and freelancers shipping to production.
min 5 users
For engineering teams that ship AI-generated code every day.
Run it
npx tryassay assess .uses: gtsbahamas/assay-action@v1lucid-mcp runs Assay inside Claude Code, Cursor, or Windsurf. Setup guide →
Run it on your own repo.
Two minutes to a verdict. No signup.
npx tryassay assess .