how many times do you have to ask?

9.00

the effective sample inside 125 answers, because the repeats bought nothing

thirty daily readings are not a sample of thirty

A language model is not a deterministic function, so one reading is one draw from a distribution nobody characterised. Refresh frequency and sample size get conflated here constantly. Asking a question once every morning for a month gives you thirty one-shot readings taken under thirty different conditions, not thirty draws of today’s answer, and the model moved underneath you somewhere in the middle of it.

what the repeats bought, as a number

In the round the front page quotes, every question came back with the same answer on all twenty asks, so the correlation inside a question was 1.000. When that happens the design effect is the cluster size, and a hundred and twenty five answers over nine questions carry an effective sample of 9.00. Exactly the number of questions. The repeats bought nothing, and the arithmetic says how much nothing.

That is not an argument against repeating. It is the reason the split is reported: the interval width is decomposed into the model answering differently and the question set you picked, so you know which of the two is worth spending the next dollar on.

the number of draws follows the rate, it is not a setting

A rate near a half needs more draws than a rate near an edge to reach the same precision, so the draw count is derived from the base rather than typed in. The floors this build uses are 2 draws at 0.3, 4 at 0.5, 6 at 0.6 and 12 at 0.75.

lulu plan --prompts 40 --brands 5 --half-width 5
lulu size --prompts 24 --runs 5 --engines perplexity,anthropic

The first sizes the round before it is bought. With no pilot it refuses to hand back a single number, because how the variance splits between the model and the question set is not something arithmetic can know in advance. The second goes the other way and prices a design you already set.

where this bites hardest

At one draw a day there is no noise floor, and without a noise floor no claim that a change moved anything can be separated from the model rerolling.

what a thin sample cannot support ± what the draws cost