Every Friday an AI agent writes the Google Ads report for a local service company we manage. On September 11 the account had its worst week since reporting began: $239.59 spent, one lead, and no phone call that lasted a full minute. The campaign had bought its cheapest clicks on record, $7.05 each. On this account, cheap clicks call less often. The agent checked tracking, the holiday, and demand, and cleared all three.
The report said so, proposed no performance change, and went to the owner with one question instead of an alarm.
Daniel Goodrich reviews the work. The owner runs the crews. The agent does the middle part. We do not name the client here, and the numbers are real.
This was also the first scheduled run under the rewritten process, where a second agent reads the draft cold and tries to break it. It broke the first draft: a grade of F, fifteen wrong numbers, and five claims the data did not support. The version the owner will read is the one that survived that pass.
This is the meta article for that run. It documents what the agent did, the calls it made, what a human changed, and what is still unknown.

The run at a glance. Red is where the hardener rejected the first draft. Gold is the investigation that explained the week.
The run in brief
The assignment was the scheduled weekly Google Ads MAA (Metrics, Analysis, Action) for the account. The report is a client deliverable for the owner. A human approves it, then posts it to the client’s project thread.
The source material was live account data pulled through the Google Ads API, read-only, for two windows: September 4 to 10, and August 12 to September 10. Before pulling a number, the agent read the client’s ledger. That file holds:
- how a lead is defined ($50 to $80 per lead is the goal, and a call counts at 60 seconds)
- every change made to the account and when
- the items on watch
- anything too new to judge
The goal was the same as every week: tell the owner plainly what the money did, and decide what, if anything, to change. This week the second half was the hard part. The obvious move on a one-lead week is to do something, and one conversion on 34 clicks is not enough evidence to do anything.
What the agent did, step by step
Stage 0: it read before it pulled
The agent read the ledger, all ten sections, then the metric spec, the shared data rules, and last week’s report. That is where it learned the counted-call unit, the target band, and the standing rule that this account’s reports leave out desktop. It also learned that a keyword re-enabled on August 28 could not yet take credit or blame for the week’s shift in queries. The two-cycle rule defers that read to September 18.
Stage 0b: it pulled the whole pack
The pack covered:
- campaign summary and budgets
- keywords with Quality Score components, across every status
- search terms
- ads with strength
- conversions by action
- network and device
- searcher location
- negatives at three levels
- call detail
- the change history
The change history returned rows for the first time after five failed attempts. The agent saved the raw pull and recorded every row count. One pull hit the tool’s output cap, so the agent replaced it with an explicit list query.
Stage 1: it drafted under the gates
The lead collapse went into Analysis and onto the watch list, not into Action. One conversion on 34 clicks does not clear the five-conversion bar. The week’s biggest-spending keyword, $65.61 on seven clicks with no lead, went to watch for the same reason.
The agent withdrew the Quality Score “reversal” that had led the previous report. Under the new two-cycle rule, a score that reads 7, then 5, then 7 while its two component directions never move is not a change. It also closed a negative-keyword gap carried for two cycles with a direct query showing all nine terms live.
Stage 2: the hardener graded it F
A second agent, in a fresh context with only the frozen draft, the raw pull, the ledger and the spec, re-derived the numbers. Fifteen did not match. One cause was self-inflicted. The first save of the raw pack carried a weekly call summary the drafter had typed by hand, and it did not sum to its own rows. Four sentences were wrong because of it.
The hardener dropped five claims outright, demoted one action to watch, and returned the draft with a punch list. One round, every verdict applied.
Stage 3: it reconciled
Every number tied to the pull. Every action either had two cycles of watch history or was a foundation break. Nothing inside the 14-day window took credit or blame. One account setting the change log could not explain went into the notes and the open questions rather than into the analysis.
Then Daniel asked the question the draft had not answered
The hardened draft said the agent could not explain the week. Daniel commissioned a same-day investigation to rule out a tracking break, a demand collapse, and noise. The investigation is its own file, and its answer changed the report.
- Nothing was broken. GA4 recorded four tap-to-call events from paid visits and Google Ads recorded four, a match through the import. The call button served on all 468 impressions. The site booked three jobs that week, all from unpaid visits.
- Labor Day was not the cause. Holidays do not suppress calls on this account. The July 4 week ran above the median on calls, and the Memorial Day week produced six tracked calls.
- Demand did not collapse. It softened about 7% while the account’s impressions fell 20%, so most of the loss was auctions the campaign did not win.
- The matching footprint was not to blame. Tightly and loosely matched traffic cost $79.37 and $79.89 a lead over seven weeks.
The finding: cheap clicks are quiet clicks. Across fourteen weeks, what the campaign paid per click predicted how often those clicks called (r = +0.65, and the relationship holds with any single week removed). This week’s $7.05 was the cheapest click price on record. Its 5.9% call rate, two calls on 34 clicks, came in slightly above what that relationship predicts. The volume matched the traffic the campaign bought, and the traffic was cheap.
Both calls were under a minute. Google counts an ad call at 30 seconds, so the account shows one conversion. We count a call at 60 seconds, so the report shows zero counted calls. That accounts for the gap between the two numbers.

Tracked calls per click by week, from the account’s call records. The three gold weeks were the outliers. The red week sits with five other weeks under 7%.
Stage 4: it appended and stopped
The agent wrote the trend row and the weekly log into the ledger, plus seven new watch items with first-seen dates, eleven open questions, and six data faults. No client view, no charts, no post. Daniel approved the draft in chat. The agent then rendered the client view and staged it for him to paste.
The judgment calls
It held the collapse off the action list
The easy report says “pause the weak keyword” or “raise the bid.” The gates said one conversion on 34 clicks cannot carry either, and the account’s own history said four of the prior thirteen weeks had come in at three tracked calls or fewer. The draft said so and proposed nothing on performance.
It withdrew its own earlier framing in the open
Since August 28 the reports had described a run of 26%, 28%, and 23% calls per click as “an elevated call rate,” as though it were the new level. It was a peak, and those three weeks were the outliers. The report says that. It also says that last week’s “back inside target” line did not survive a week.
It proposed a watch trigger that would not cry wolf
The first draft proposed escalating if the next week produced fewer than three counted calls. Checked against history, that trigger would have fired on four of the last thirteen weeks. The proposed replacement sits in the ledger pending Daniel’s approval. It tests the mechanism instead of the symptom:
- If click price recovers to $11 and the call rate stays under 10%, the relationship has broken and we escalate.
- If price stays under $9 and calls stay low, the cheap-auction pocket is persisting and the bidding question opens.
- If both recover, the item drops.
It narrowed the ask to the owner
The first draft asked the owner to help rule out a tracking break. After the investigation ruled it out from our side, the ask became a two-minute confirmation. The agent asked whether two short calls, one Saturday morning and one Wednesday afternoon, matched what the owner saw on the phones.
It logged an unexplained change as a question
Someone had changed one account setting with no record of who or when. The report records that as an unlogged change and an open question. It does not build a story on it.
What it cost
We did not measure this run. Our July run articles estimated about 14 agent-minutes and $2 to $4 per run, against an hour to an hour and a half of manual work for a draft with no hardener pass. This run did more, since the hardener pass and the investigation are new, and we record its time, tokens, and cost as UNKNOWN. The notes block now requires start and end timestamps, so the next cycle will carry numbers.
Where the humans were
The agent did these on its own:
- read the ledger
- pull and save the data
- write the draft under the gates
- dispatch the hardener and apply its verdicts
- reconcile
- append the ledger and trend file
- render the client view once approved
- write this meta article
Required a human this run:
- commissioning the investigation, since the hardened draft had stopped at “we cannot explain it,” and the investigation changed its central finding
- approving the draft
- pasting the client view into the client’s project thread
- approving the replacement watch trigger
- the answer only the owner has
Two of those touched the report itself, and one of the two changed the analysis. The hardener handled every number correction. The back-test of the previous week, run through the same process a few days earlier, had needed zero. The gap between those two counts is what we work to close.
What this run changed in the recipe
- Compute derived blocks from rows. Never type them. A hand-typed weekly summary in the raw pack produced four wrong sentences. The rule is now in the analyzer, the reviewer’s re-derivation step, and the query pack.
- Paste the hardener’s output verbatim. This report described the verdicts in one sentence of counts. Only one of the four reports that cycle printed the verdict table, and none printed a grade after the round. The skill’s completion checklist now requires the table, the counts, and both grades.
- Test a watch trigger against history before proposing it. “Fewer than three calls” fires four times in thirteen weeks on this account. Check the false-positive rate on the series before writing the condition.
- Treat “we cannot explain it” as a prompt for a pull. The explanation was in data the standard pack already had: weekly cost per click and call rows. The next revision of the recipe should ask, before closing Analysis, whether the draft has tested any mechanism in the pack against the symptom.
- Name the levels when claiming nothing is disapproved. This account’s report did not make that error this week. The same cycle on another account said “nothing is disapproved” from an ad-level pull while two image assets sat disapproved. The rule now applies everywhere.
Run inventory:
- Execution ID: GAA-2026-09-11-C1A
- Recipe: google-ads-analyzer v1.5.0, before the house-format revision, hardened by google-ads-reviewer v1.5.0 in a fresh context, one round
- Windows: 7D 2026-09-04 to 09-10, and 30D 2026-08-12 to 09-10
- Hardener result: grade F on the first draft, 15 numbers corrected, 5 claims dropped, 1 action demoted to watch
- Actions in the final report: 5 (2 for the owner, 3 for the agency), and nothing shipped to the account this cycle
- Ledger appended: 1 weekly-log row, 7 watch items from this run, 11 open questions, 6 data faults, quarantine updated
- Human interventions on the report: 2 (Daniel commissioned the investigation, and Daniel approved the draft with the reframed ask to the owner)
- Time, tokens, cost: UNKNOWN
- Evidence, in the private client folder:
deliverables/2026-09-11-google-ads-maa.md(Generation notes block),deliverables/2026-09-11-call-volume-investigation.md,compiled/ledger.mdrows dated 2026-09-11, andcompiled/trend.csvlast row
Run it on your own account
Install the Google Ads MAA skills from GitHub, onboard one account, and let the scheduled run fire on Friday. Read the hardener’s verdicts before you read the draft. The Google Ads MAA Agent recipe describes the five stages this run followed and the setup steps.
Two earlier runs under the previous process show what changed: the July review that fixed a blind spot and the July painting contractor run. Dennis Yu’s MAA framework is the discipline underneath all three.
This meta article documents one client’s Google Ads MAA run on 2026-09-11, execution GAA-2026-09-11-C1A. The client is anonymized, and the numbers are unchanged.
Agent receipt: localservicespotlight.com draft · Claude [claude-fable-5-1] · action: drafted · human review: Daniel Goodrich reviews before publishing
Originally published .
