Warning: Most people testing AI image tools are judging them all wrong
I spent 3 hours last Saturday at my desk in Austin running the same 5 prompts through 4 different image models, and I noticed something weird. Everyone in the comments keeps ranking these tools by how pretty the output looks, but the real gap is how well they follow tiny instructions like 'blue mug on the left, red book behind it.' I ran 20 test rounds and the prettiest model got the object placement wrong 14 times, while the boring looking one nailed it 18 times. If you build anything client facing, count how many tries it takes to get the layout right instead of just saving the pretty ones. Has anyone else tracked instruction following this way, or am I overthinking the benchmark?