Protocol
How we
test
One room, one rubric, 35 units and 2,110 logged hours. Published in full so you can decide whether you agree with the weights.
- Test room
- 18 x 14 ft, 9 ft ceiling
- Conditions
- Door closed, HVAC off
- Units per brand
- 3 to 4
- Session length
- 4 hours maximum
On the bench
Eight seconds of a controlled burn
A single candle burning in the Chicago Candle Factory test room. Every score in the index comes from sessions like this one: a fixed room, a closed door, and a flame left alone until the melt pool reaches the glass.

Step by step
The burn protocol
- 01
Buy it like a customer
Every unit is bought at full retail from a normal retail channel, in the size most people actually buy rather than the size that tests best. Three to four units of each brand, so a single bad pour cannot set a score.
- 02
Record the cold throw first
Lid off, unlit, at arm's length, scored before the candle is ever lit. Doing this first matters: once you know how something performs hot, you cannot un-know it when you go back to score the cold throw.
- 03
Weigh, measure and photograph
Vessel diameter, fill weight and wick count go in the log. Diameter sets the length of the first burn, at one hour per inch, which is the single biggest factor in whether a candle tunnels.
- 04
Run the first burn to a full melt pool
In the test room: 18 by 14 feet, 9 foot ceiling, interior door closed, HVAC off, no other fragrance in the space for 12 hours beforehand. Time from first light to the melt pool reaching the vessel wall is recorded to the minute.
- 05
Walk the line at two hours
A tester walks a marked line away from the candle and records the last distance at which the scent is unambiguously identifiable. Repeated across every unit and averaged. This is the throw distance figure on every review.
- 06
Burn to the residue floor
Sessions of no more than four hours until the candle will not sustain a flame. Tunneling, soot, wick behaviour and flame height are logged each session, and the residue depth is measured at the end.
- 07
Score against the fixed rubric
Six criteria out of 10, multiplied by fixed weights, summed to a lab score out of 100. The score is calculated from the log, not written to justify a conclusion.
Weights
The rubric
Hot throw carries the most weight because it is the only thing a candle does that a nice jar cannot fake. Sustainability carries the least, not because it matters least, but because it is the criterion we can verify least independently.
- Hot throw 25%
- Scent measured at distance in an 18 by 14 foot room, door closed.
- Burn quality 20%
- Full melt pool, tunneling, soot, wick behaviour, end residue.
- Cold throw 15%
- Scent from an unlit candle at arm's length, lid removed.
- Value 15%
- Performance and residual value per dollar, not raw cheapness.
- Design 15%
- Vessel, label, and whether the object survives the wax.
- Sustainability 10%
- Wax source, packaging waste, and what happens to the vessel.
Weights sum to 100. Bars are drawn at 4x for legibility.
Limits
What we do not do
- We do not accept review units, gifted product or press samples.
- We do not run affiliate links, so no ranking position pays us anything.
- We do not test on a brand's schedule or around a product launch.
- We do not score a brand on a scent we happen to dislike. Composition preference is noted in the copy and kept out of the numbers.
- We do not use lab instruments we cannot afford to calibrate. Throw distance is a human reading and we say so.
Honesty
Where this is imperfect
Throw distance is a human reading. Noses fatigue, and the same tester on two different days will differ by a foot or two. We average across units and we keep the room, the table and the marked line identical, which makes the numbers good for comparison between brands and only roughly right in absolute terms.
Scent preference is not measurable and we do not pretend otherwise. If a composition is not to our taste, that goes in the copy where you can weigh it yourself. It never touches the score.
Sustainability is the weakest column. We can verify a wax type when a brand publishes it, we can weigh packaging and we can judge whether a vessel survives the wax. We cannot audit a supply chain, so brands that disclose nothing score lower than brands that disclose fully, which is a proxy rather than a measurement.