How honest is Dress Aloud?
Updated 16 September 2026. This page is updated after every measurement, including the ones that go badly.
The problem we measure
A language model asked "does this fit" will happily invent a sleeve length or a "typical medium" chest measurement. A sighted shopper spots that. A blind shopper cannot. So Dress Aloud's fit verdict is written only from facts the model was given: your measurements, the size you are considering, and the store page's title, fabric, fit words and size chart.
The guard
A number guard, ordinary code rather than AI, checks every sentence of the verdict before you hear it. Any number that is not in those facts, and is not the difference between one of your measurements and a size-chart number, is sent back once with the invented numbers named. If the rewrite still invents, that sentence is dropped and you are told a sentence was dropped and why. Praise words are forbidden and counted afterwards.
What we measured
The same 18 try-ons each time: two synthetic subjects, real product pages on Amazon, Uniqlo and Zappos, tops and shoes, a size chart supplied for half of them, several sizes wrong on purpose. An independent script, stricter than the guard, re-checks every delivered verdict against the facts.
| What we count | First guard | Second guard | Third guard, the one live now |
|---|---|---|---|
| First drafts that invented a number, all caught and rewritten | 7 of 18 | 5 of 18 | 0 of 18 |
| Delivered verdicts still containing an invented number, strict check | 3 of 18 | 1 of 18 | 0 of 18 |
| Delivered sentences with an invented number | 4 of 108 | 1 of 113 | 0 of 103 |
| Sentences dropped after the rewrite | 0 | 0 | 0 |
| Praise words delivered | 0 of 18 | 0 of 18 | 0 of 18 |
| Median time to the spoken verdict | 45 s | 47 s | 36 s, then 31 s after a speed change |
The misses, in plain words
- First guard. Three verdicts on pages with no size chart asserted a "typical" garment measurement as fact, for example "a standard medium measures at least 44 inches". The guard let them through because its arithmetic rule accepted the difference between any two facts, so a height minus a head circumference could justify almost any number. Closed the same day.
- Second guard. One strict hit: the number 100 in "100 percent cotton", read straight from the store page. The test harness had not saved the fabric line for outfit pieces, so the checker could not see it. The app's own guard, which had the page, was right to let it through. Read strictly it is 1 of 18; with the page in hand it is 0 of 18. Both are stated because the strict script is the one that cannot be argued with.
- Third guard. Nothing invented and nothing rewritten on the 18 items, on two separate runs. A later run with the same result also found that a shopper's typed note had been silently dropped from the verdict whenever a store page was read. That was a bug in our code, not the model's, and it is fixed.
What these numbers are not
They are the guard's interception rate on synthetic subjects, 18 items. "0 of 18" means 18 clean items on one run, not a percentage, and not an accuracy rate. The measurement that matters is people: blind listeners deciding from the spoken verdict, sighted raters scoring every claim against the picture, and shoppers telling us whether what they bought fit. That study is written and starting with our first testers. Its results go on this page, misses included.
What you can do
If a verdict is ever wrong about you, say so: the thumbs-down under every result, the "Report this picture" control, and the follow-up question the app asks a week after a purchase all reach a person. That is the data that keeps this page honest.