Does it work beyond the example?

A system gets 99% of its training photos right but fails at night. Which test shows this?

The team wants to know if it works outside what it has already seen.

Getting the known example right is only the start. 🔎