KEY POINTS
Key points
- When Search Console impressions fell from more than 700 a day to 2, no reasonable range of normal variation could contain the drop. It should have been treated as a warning sign, not a reporting delay.
- The cause was in the Page indexing report: 687 pages were marked “Crawled – currently not indexed”, most of them thin glossary pages generated in a single batch.
- Following the Drivetrain Model (objective, levers, data, models), the thing to watch was indexing, the lever, not impressions, the outcome.
- Much of the GA4 traffic before launch came from the team itself. Internal traffic needs to be filtered out before launch, or external visitors and the team’s own activity cannot be told apart once promotion starts.
At the end of August I started a course at Minerva, CS312, which is roughly about making decisions based on data.
After the first class introduced a framework called the Drivetrain Model, the pre-class materials ended with a simple self-check:
Think of a problem you have faced at work and describe it again using this framework.
I tried a few examples. None of them felt right, so I put the question aside.
Then this week, as the team was getting a Data Studio templates and tutorials website ready to go public, I went back through the last three months of data, layer by layer from Search Console to GA4. That was when I realised:
the answer to that question had been in front of me the whole time.
That day in July
The website had not been officially launched.
But so that we could build, test and adjust as we went, it had been live since the summer, and tutorials were being added one by one.
By early July it was already getting impressions in Google Search.
On the best day, more than 700.
The next day, 2. From then until early October, it was almost always 0.
The team discussed it in a meeting at the time.

Our view then was that GA4 still showed people arriving from organic search, so the site had probably not disappeared from Google. More likely, Search Console had not reported its data in full, and we could keep an eye on it for a few more days.
There was also something nobody said out loud, which in hindsight was perfectly understandable:
the website wasn’t officially live anyway.
So the matter was set aside.
That explanation wasn’t unreasonable. Data tools do sometimes lag, miss data or fall out of sync. None of that is rare.
The real problem was this:
it was the first explanation that sounded reasonable, and we stopped there.
It was only when we checked again in October that we found the actual cause.
It wasn’t a bug in the tool. Before the site had even launched, Google had already made a judgement about it and left a large number of its pages out of the search results.
Before explaining a change, ask whether it is normal
One CS312 class spent a good deal of time on prediction intervals.
Put simply, if you had to guess how many drinks a bubble tea shop will sell tomorrow, a better answer than “120 cups” might be:
tomorrow will probably be somewhere between 100 and 140 cups.
What struck me about this idea wasn’t the forecasting itself. It was that it can be turned around to help answer a different question:
is what happened today still normal?
If the shop sold 118 cups today, that is within the range and probably just ordinary variation.
If it sold 3, “business was a bit slow today” doesn’t cover it. Something has changed.
The same goes for the website data in July.
For a site still being prepared, a few more impressions one day and a few fewer the next is normal. But a fall from more than 700 to 2 is hard to fit inside any range of normal variation, however wide you make it.
At that point the first reaction shouldn’t have been “maybe the tool is acting up”, but:
“what has changed?”
Looking back, what was missing then wasn’t a report. It was this one step of judgement: first decide whether this is normal variation or a warning sign, then decide whether it needs following up.
Going through it again with the Drivetrain Model
The Drivetrain Model isn’t complicated. It simply reminds you, before diving into a pile of numbers, to answer four questions in order:
1. Objective: what result do we actually want?
The real goal of this website isn’t “more traffic”. It is:
to help people looking for tutorials on Data Studio, GA4, Search Console and related tools find this content through search.
2. Levers: what can we actually control?
Impressions aren’t something you can adjust directly. What you can change is:
- which pages Google is able to index
- whether the content of a page is worth indexing
- whether the site structure and content quality make sense to a search engine
In other words, rather than watching how far impressions have fallen, the thing to look at is the lever further upstream.
3. Data: which data best tells us whether the lever is working?
The answer isn’t just the impressions and clicks on the Search Console overview. It is the Page indexing report.
That report goes page by page: which pages are indexed, which aren’t, and why not.
The answer had been sitting there since July. 687 pages showed:
“Crawled – currently not indexed”.
In other words, Google had visited them but decided, for now, not to include them in its search results.
Digging further, we found that most of these were glossary pages for individual metrics, generated in one batch at the time to build up content quickly. Each page had very little content, and the pages were very similar in structure. We had thought we were adding to the site’s content. In the end, they may well have weakened its overall quality signals.
The data had never gone missing. It was just that, at the time, we were looking at the outcome at the end rather than the lever before it.
If Search Console is new to you, the site has a beginner’s guide (in Chinese): Search Console basics: how to read impressions, clicks, CTR and average position.
4. Models: how do these levers affect the final result?
This step is fairly intuitive.
The more pages that are properly indexed and have real content value, the better the chance of being found in search. Conversely, a site full of similar, very thin pages won’t necessarily perform better in search just because it has more pages.
More isn’t necessarily better.
Often, what really needs managing isn’t quantity but the quality of the signal.
Someone hiding in the numbers
The check turned up something else I found interesting.
From August to mid-September, GA4 showed steady traffic: between a dozen and thirty-odd sessions a day on weekdays. Looking only at the totals, you might even think: not bad, people are coming before we have even launched.
Once the data was broken down, the picture started to look different.
The traffic appeared almost only on weekdays and was close to zero at weekends. More than half of the engagement came from a single district. And quite a few sessions lasted unusually long.
By then the answer was fairly clear.
These weren’t enthusiastic early users.
It was us.
During the preparation period, the people visiting the site most often were the colleagues uploading templates, writing articles, editing content and testing features. The team’s working rhythm was deciding two things at once: how much traffic the site had, and how “active” it looked.
When CS312 covered “correlation is not causation”, it used a familiar example:
on rainy days, convenience stores sell more umbrellas, and there are more puddles in the street. But selling more umbrellas doesn’t make the puddles deeper. There is a common factor behind both: the rain.
In analysis, a factor like this, left out of the model but affecting what we observe, can be thought of as an omitted variable.
For this website, the “rain” was:
internal traffic.
If this isn’t sorted out before launch, then once promotion begins, external visitors and the team’s own activity are all mixed together. When traffic rises, it is easy to conclude that the marketing has started working.
But is it really new visitors, or has the team simply been busier this week? That becomes very hard to answer.
GA4 has its own way to identify and filter out internal traffic, and it isn’t technically difficult. The one thing to be careful about is that once a data filter is active, the excluded data cannot be recovered, so it is best to leave it in “Testing” for a few days first.
What tends to happen during preparation is the feeling that:
“we can deal with this later.”
But many data problems grow quietly inside “later”.
For more on reading common GA4 data anomalies, see this guide (in Chinese): GA4 (not set) and Unassigned: four common data anomalies explained.
Frameworks are neat; real life usually isn’t
At this point I want to say a word in defence of the judgement we made that day in July.
That meeting had more than a dozen items on the agenda, and the Search Console anomaly was just one of them. Reporting delays do happen, and the site wasn’t officially live.
Put back in that moment, rather than looking at the answer three months later, the judgement wasn’t that unreasonable.
That is also why I have come to like these frameworks more and more. They aren’t there to prove “see, you got it wrong”. They help build a steadier order of thinking.
The Drivetrain Model won’t hand you the answer, and a prediction interval won’t suddenly flag what is broken. What they offer may simply be a few questions that slow you down a beat:
Is this change normal?
Am I looking at the outcome, or at the thing that can change the outcome?
Has someone, or something, been mixed into this data?
When things are busy, these are the questions most easily skipped. Yet the most valuable judgements are often hidden in that one extra question.
So lately I have been reminding myself not to wait until something goes wrong before reaching for a framework, but to make it a regular habit in meetings, when reading reports and when making decisions.
Much of what is taught in class isn’t a grand new idea. The hard part is remembering to use it when everything is rushed, the information is incomplete and ten other things are waiting.
Before launch is the cheapest time to read the signals
Since this check, the team has started removing thin pages, reorganising the site’s content and resubmitting key tutorials to Google. Internal traffic in GA4 will also be handled separately, so that the data after launch is closer to how real users behave.
Doing all this now is still relatively cheap.
The site hasn’t launched, no large promotion budget has gone in yet, and nobody has to fix the data while a campaign is running.
Discover it only after launch, and the cost is more than a few hours of changing settings. It can also include wrong judgements, optimising in the wrong direction, and missing signals that could have been seen much earlier.
So that CS312 self-check took me about two months to hand in.
What I have kept from it isn’t just the four steps of the Drivetrain Model, but a question I now ask myself often:
having seen a number, do I really understand it, or have I just found an explanation that sounds reasonable?
And one more question for anyone getting a new website, product or service ready:
on launch day, which number will you look at?
And:
before launch, had that number already been trying to tell you something?
Frequently asked questions
Is a sudden drop to zero in Search Console impressions a bug?
It can be, but it is better not to assume so. Search Console data is sometimes delayed, but if impressions fall far below their usual range and stay there for several days, the website itself has more likely changed. Start with the Page indexing report and check whether a large number of pages have been left out of the index.
What does “Crawled – currently not indexed” mean?
Google has crawled the page but has not added it to its search index yet. It may be indexed later, or it may not. If many pages are in this state at once, it is worth checking whether they have too little content or are too similar to each other.
What are the four steps of the Drivetrain Model?
In order: objective, levers, data and models. First define the result you want, then identify the factors you can actually control, then work out which data shows whether those factors are having an effect, and only then build a model of how they affect the result. The approach was introduced by Jeremy Howard, Margit Zwemer and Mike Loukides in 2012.
Should GA4 exclude internal traffic before a website launches?
Yes, and the earlier the better. Before launch, much of the traffic comes from the team itself. If it isn’t separated, external visitors and internal activity will be mixed together once promotion starts, making it hard to tell whether the promotion is working. A GA4 data filter cannot be undone once active, so it is best to check the scope in “Testing” first.
Further reading
The website in this story is CloudAD’s Data Studio templates and tutorials site, with hands-on guides to GA4, Search Console and Data Studio (in Chinese):
- Data Studio templates and tutorials: all tutorials
- Search Console basics: how to read impressions, clicks, CTR and average position
- GA4 (not set) and Unassigned: four common data anomalies explained
- Can you trust what AI tells you? Research finds it agrees with your mistakes 63.7% of the time (COO perspectives: why we tend to stop at the first reasonable explanation)
References
- Google Search Console Help, Page indexing report
- Google Analytics Help, Filter out internal traffic
- Jeremy Howard, Margit Zwemer and Mike Loukides, Designing great data products, O’Reilly Radar, March 2012
- Rob J Hyndman and George Athanasopoulos, Forecasting: Principles and Practice (3rd ed.), 5.5 Distributional forecasts and prediction intervals



