| |
Your model already knows the answer: how benchmark answers leak into LLMs
Large language models frequently achieve high benchmark scores not through genuine reasoning but by having encountered the answers during training, a problem called "data contamination" that affects many widely-used tests. The article identifies three routes through which answers leak into models: input leak (reading documents that reveal outcomes at test time), benchmark leak (training data containing the benchmark itself), and outcome leak (absorbing publicly reported results during training). This fundamental challenge undermines the validity of benchmarks built from real-world events with public outcomes, making it difficult to distinguish between models that truly reason versus those simply recalling memorized information.
Read Full Article →
← More Tech news