Essays

The audit grades evidence, not craft

At 3:31 on a Monday morning in September, Google sent me an email that only arrives once an app has made it through Google's own review. Dustav had passed, and now it needed a security assessment called CASA, due by December 13. Nine days later the lab's report came back: 48 requirements, 48 pass, zero fail. The Letter of Validation followed that week, and on September 29 Google approved the app. Anyone with a Gmail address can now connect Dustav without clicking through a scary warning screen.

I build Dustav alone, with Claude Code doing the typing. The online folklore about this process is mostly horror stories: months in a queue, five-figure audits, apps rejected for reasons nobody explains. Ours took two weeks of calendar time and $675. I don't think that's because the code is special. I think it's because of something I got told early and didn't quite believe:

The audit doesn't grade your craft. It grades your evidence.

Two gates, not one

If an app reads people's Gmail, Google calls those permissions restricted scopes, and restricted scopes come with two separate checks. People blur them together, and it's worth pulling them apart, because they fail for different reasons.

The first is Google's own review, and it's opinionated. A person at Google checks whether your app does what you say it does. You declare every permission by hand. The console never fills them in from what your code actually requests; they're two separate lists, and it's on you to make them match. For each permission you write a short justification, and you record a video showing each one being used exactly as described. The reviewer watches the video with your text beside it.

The second is CASA, and it's technical. An outside lab assesses you against a slice of the OWASP security standard: authentication, sessions, access control, encryption, input handling, logging. Google tells you it's required, sets a deadline, and points you at a list of approved labs.

Google's gate fails you for saying something untrue. CASA fails you for being unable to prove something true. Everything below is one of those two sentences.

Declared equals deployed

The rule that got us through Google's review in two days was one we'd adopted weeks before for a different reason. The code is the source of truth, and every document conforms to it.

That sounds obvious until you watch it go wrong. A month before submitting, I asked whether Google's review was purely technical or whether they could be opinionated. The answer was that the opinionated part is exactly this: a reviewer comparing what you claim against what they can see. So we audited every document Google would read, line by line, against the deployed app, and brought each one up to date with what the code actually does. The direction matters. The tempting move, when a document and the code disagree, is to change the code to match the document. That's the most dangerous direction there is. Code that conforms to a stale document is a bug with paperwork. So the rule is written down: never change behavior to match a document; change the document.

It kept earning its keep. We dropped a Gmail permission we'd been requesting but didn't need, because it was a strict subset of another one we already had. And every justification describes exactly what the video shows. For calendar invites, that's: Dustav prepares the invite, and the owner's tap sends it.

The video is a proof

The standard advice is that the demo video is the most common reason apps get rejected, and every cause on the list is mechanical. The URL bar isn't visible during sign-in. A permission you declared never appears on screen. There's no English narration. You recorded against a development server.

So I treated the video as a proof, not a demo. It was eight minutes, one take, against production, on a demo account with fictional mail. It goes one permission at a time. For each thing Dustav does, I cut to a real Gmail tab and show the effect: the drafted reply sitting in Gmail's Drafts, threaded into the conversation. The filed email wearing its label. The deleted one in Trash, recoverable. The sent reply in Sent, and gone from Drafts. A calendar event appearing in Google Calendar only after the tap, then moved in Google and reflected back in Dustav. A reviewer checking my text against the video never has to trust me, because every claim has a frame.

Google's stated timeline was four to six weeks. It came back in two days, with six of seven stages green on the first pass. The seventh stage was the lab.

The lockfile was the real audit

In August, weeks before any of this, I'd said something breezy: the app is well built, so CASA should be fine. The calibrated answer I got back is the thesis of this essay. The assessment doesn't look at how thoughtful your architecture is. It looks at evidence, and part of the evidence is a machine reading your list of dependencies.

That check takes thirty seconds, and it's blunt on purpose. A scanner can't see how you've configured a library. It reads version numbers and files a finding. So "it's not reachable the way we use it" is an argument, and an argument isn't evidence. The week before the lab's scan (not earlier, since advisories publish weekly), we upgraded our way to a dependency check that reads zero.

The best lesson came out of that cleanup: a safety setting a library applies for you is a setting you don't control. An upgrade can change a default without a word. So any protection we rely on is set in our own code, on every call, and a test pins it there. If a future upgrade ever changes the default, nothing changes for Dustav, and if someone removes our line, the build goes red.

Forty-eight questions, answered from the code

The questionnaire is 48 security requirements, and every answer needs evidence attached. The portal takes images only, so PDFs were refused.

We didn't answer from memory. Three read-only audits walked the codebase: one on authentication and sessions, one on access control and permissions, one on input handling and configuration. Every claim was tied to the file and line that makes it true. A script rendered each requirement into pages: the statement, how it was checked, and the cited source quoted verbatim with line numbers and the commit it came from. Then those pages were rasterized into images: 91 of them for 48 answers.

We wrote one rule at the top of that answer sheet: where the truthful answer is "not applicable" or "partial," it says so, with what compensates. An over-claim costs you a rescan. An honest note costs nothing. One answer went in as partial. That's the whole trick, and it's not a clever one.

Round two: nobody disputed a control

Thirty-eight of the 48 passed first time. Ten came back with tester comments, and the pattern in those ten is the most useful thing I learned all month:

The assessor didn't dispute a single control. Every comment asked for one piece of evidence the first answer didn't include. Show the code that creates a fresh session at login. List every endpoint that takes an ID from the user. Date the TLS scan. Say how each key is generated and rotated. Show a real log line from a real payment.

All ten were answered the same day, and a few of them left the product a little better documented: a line in the deploy runbook, a written secrets policy, a richer audit log. The payment log sample had been rotated out of existence, so I topped up my own account by $50 to produce one fresh line of evidence. Where a first-round answer had been loose, the reply tightened it in the open and moved on.

Then, a few days later: 48 pass, zero fail.

Why it went this way

I'd love to say we passed because the code is exceptional. The real reasons are more boring, which is why I trust them.

The product's shape made the answers short. Dustav is single-player: one person, one agent, and every route checks that the thing you asked for belongs to you. So "is access control enforced?" was answered with an inventory of about 120 routes, each with the one ownership check it goes through. The answer had one shape, not dozens. Nothing sends by itself, so the Gmail justification could make a flat promise instead of a hedged one. Gmail stays the store, so there was no parallel copy of anyone's mailbox to explain.

The evidence already existed. This is the part I didn't expect. Dustav has a hundred-odd verification scripts that pin its behavior, a folder of design documents written before the features they describe, and a session log recording why every decision was made. None of that was built for an audit. It was built because a solo founder building with an AI needs a memory that outlives any one session. But an audit is a request for evidence, and it turns out we'd been producing evidence as a byproduct for four months. When the assessor asked for an endpoint inventory, it was an afternoon's sweep, not a project.

The work split cleanly. Claude read the code, drafted the answers, cited the lines, and rendered the evidence. I did the parts only a person can do: signing in on the real server, confirming two-step verification on the admin mailbox, taking the dated screenshot, spending the $50. The same collaboration that wrote the code answered for it. That's only possible when the code was written to be answerable.

The rule

It isn't over. CASA recertifies every year, so this happens again next September. I'll do the dependency sweep the week before, not the month before. And as I wrote in August, this is the rent on the best real estate in software: the place your customer already lives. Price it in.

But here's what I'd tell anyone facing it for the first time. Security is a claim. An audit is someone asking you to prove it. If your code is the source of truth, your documents match it, and your tests pin what you promise, then the proof is a query you run, not a project you start. Build so the evidence already exists. Then the audit is paperwork.