System and method for prediction of artificial intelligence model generalizability for unseen data

Problem

Artificial intelligence models often perform well during development but show unpredictable drops in accuracy and reliability when deployed on data that differ from their training sets. In high‑risk settings such as clinical care, these shifts can arise from changes in hardware, protocols, institutions, time, or patient populations. Most existing strategies attempt to improve generalizability indirectly by increasing training data size, or applying data augmentation, yet these approaches rarely address real‑world variation or indicate when a model should not be trusted. As a result, AI systems may produce confident yet unreliable predictions on out‑of‑distribution data, creating patient safety risks, regulatory challenges, and barriers to widespread adoption of AI in critical settings. There is a need for deployable AI methods that can detect when models are operating outside their domain of validity and warn users before unreliable predictions occur.

Solution

Researchers at The Ohio State University have developed a technology that enables AI models to assess and report their own generalizability on a case‑by‑case basis by incorporating a statistically grounded self‑evaluation mechanism. During training, a latent space mapping (LSM) approach forces the model to organize its internal feature representations into a well‑defined multivariate normal distribution. When new, unseen data are presented, the system measures how closely those data align with the distribution learned during training using established distance metrics. If the new input deviates significantly from the training distribution, the model identifies the case as having low generalizability and can generate a warning or flag the output as low confidence. Rather than attempting to guarantee that a model will perform equally well across all possible scenarios, the approach provides a practical and transparent way for AI systems to determine when their predictions are likely to be reliable and when additional human review or alternative workflows are warranted. The approach was tested in the context of a brain metastases detector for T1‑weighted contrast‑enhanced 3D MRI, but it is applicable for most classification deep neural networks (DNNs).

Applications

  • AI for medical imaging
  • Clinical decision support
  • Autonomous and Safety-critical systems (robots, self‑driving vehicles, and automated industrial equipment)

Advantages

  • Model‑agnostic: works with existing AI architectures without redesign
  • Provides on‑the‑fly confidence assessment, warnings and decision gating
  • Statistically grounded: uses well-known statistical metrics instead of heuristic scores
  • Clinically validated: demonstrated on multi‑institution MRI datasets

Seeking opportunities for out-licensing and collaboration

Loading icon