Three AI reviews of the same project, and where they disagreed
I run a YouTube Shorts channel that is almost entirely automated. It publishes four short factual videos a week, and back in June it got dropped out of the Shorts feed overnight, which I have written about elsewhere. Since then the numbers have mostly recovered. Average view percentage went from around 55% to 84%, median views are back over a thousand, and several videos are now looping — watched past 100% of their length.
None of that is producing subscribers. In the last seven days the channel added zero, at roughly nine hundred median views a video. Over its whole life it converts at about 0.12%, where a healthy Shorts baseline is somewhere between 0.5% and 1%. So the machine is working and the channel is not, and I had stopped being able to see why.
What I did about it was give three different AI models the same set of channel data and ask each of them, separately, to tell me why it was not growing and what they would do. None of them saw the others’ briefs or answers. Then I gave a fourth model all three reviews and asked it to reconcile them.
The unanimous part was the cheap part
All three agreed on five things. Do not restore the second daily upload slot that has been paused since July. Freeze the weekly long-form compilation, which has no audience and no search intent. The sameness of the template is the dominant risk to the channel, not just as a policy matter but as the reason views are not converting. The 0.12% is a product problem rather than a distribution one. And the original justification for that second upload slot had already been settled against by an experiment months ago, so the resume condition I was still tracking was measuring something that no longer mattered.
I acted on all of it, but I would not want to overstate what the agreement is worth. Three models converging is not three independent opinions — they are trained on overlapping material and they share a house style of advice about audience growth. Where all three land in the same place I take it as evidence that the answer is conventional, not that it is correct. The useful thing about that list is that everything on it costs nothing but restraint, so being conventional and being right are hard to tell apart and it does not matter much which it was.
The split was about severity, not facts
Nobody disagreed about the numbers. All three read the same figures and produced the same arithmetic from them. The disagreement was about how bad it is.
The first treated it as a tuning exercise: rotate three or four opening patterns, add a subscribe card, rebalance the subject mix, and expect it to come good inside ninety days. The second wanted the variance built into generation rather than added by hand, and was mainly worried about the template being fingerprintable. The third refused the framing entirely and said the thing is not under-optimised, it is a commodity — that reducing the upload rate on an unchanged product was, in its words, quieter spam, and still spam.
That third review is the only one that changed what I think. Its evidence was sub velocity against retention: before the cliff, ninety-four subscribers across a hundred and twenty-six videos; after it, seventeen across sixty-four; last week, none at all while retention was at its best-ever level. People finish the video and leave. At 0.12%, getting to a thousand subscribers needs the better part of a million more views, and that is the sort of number that tells you the strategy is not slow, it is wrong.
The harshest reviewer had the best diagnosis and the worst remedy
Having got there, the same review then told me to undo the niche narrowing — bring back the subject areas I had zeroed out, several of which had better median views than the ones I kept. The argument is reasonable on its face and the medians are real.
It is also a recommendation to abandon a deliberate thirty-day test at day fourteen. I narrowed the subjects on purpose, to find out whether clustering the channel on one theme does anything for it, and stopping halfway leaves me with no reading at all rather than a negative one. I kept the test and logged the objection as a controlled experiment for day thirty-one.
There is a pattern in that worth naming. The reviewer most likely to identify the real problem was not the one whose remedy I should take, and I do not think that is a coincidence — a model prompted to be brutal is being asked for a disposition, and willingness to say the uncomfortable thing travels with willingness to say the drastic thing. Only one of those two is calibrated by evidence.
A recommendation resting on four videos
The clearest failure was quieter than any of that. One of the reviews recommended raising the weight on two subject areas, with sound reasoning about how the audience was clustering. One of those subjects had two videos behind it. The other had four.
Nothing in the output marked that out. It arrived in the same register, at the same length and with the same confidence as the recommendations resting on a hundred and twenty-six videos, in a document long enough that I was not going to stop and check the sample behind every line unless I had decided in advance to do exactly that.
I should say that this is not a machine failing at something people are good at. An earlier weighting pass on the same channel was mine, and it ranked subjects on average views across a hundred and sixty-one videos. Averages skew badly on a couple of viral outliers, and when I re-ranked on medians of the videos from before the collapse, one subject I had weighted near the top turned out to be third from the bottom. The error is ordinary. What is different is that a model produces it in fluent, confident prose at volume, so the ordinary defence of noticing that a claim feels thin does not fire.
What the fourth model was good for
Reconciling three long reviews is genuinely useful and it is clerical work, which is exactly what I would hand to a model. It listed the agreements, mapped the four axes of disagreement, and turned the survivors into a sequence I could actually run.
It did better than I expected on the sample-size problem, holding the two re-weights back on the grounds that two and four observations are not enough. It did about what I expected on the substantive split: given one review saying the product is broken and two saying it needs tuning, it took the diagnosis from the first and the remedy from the other two. That is a defensible call and it is also the only move a synthesiser has. It cannot decide that one reviewer is right and two are wrong, because it has nothing to weigh them with other than the reviews themselves.
What shipped
The long-form is frozen. There is now a rule preventing the same subject area from recurring inside three uploads, which I checked by running five hundred selections against the real subject list and the real history rather than trusting the deployment — no violations, and no repeat inside the window across a simulated run. That one was overdue on its own evidence: the three uploads before it were ancient engineering, ancient engineering and metallurgy. And the closing line of each script now has to loop grammatically back into the opening line, which is the one idea from the review I would not have arrived at.
Two things came out of it that a script cannot do. The first eight to twelve words of every script get rewritten by hand now, and I pick the opening frame myself. Not because a model could not do either, but because the whole diagnosis was that the output is too uniform for anyone to want more of it, and automating the variance produces uniform variance.
The exercise cost an evening. Most of what came back I could have written myself on a good day, and the part I could not was the argument between them.