When models disagree, it’s a red flag that an input might be risky or unusual....
https://emiliosbestinsights.rivetgarden.com/posts/how-to-build-a-small-labeled-dataset-from-disagreements
When models disagree, it’s a red flag that an input might be risky or unusual. Measuring ensemble variance or entropy helps spot these tricky cases