GAN evaluation metrics, for engineers | What FID and Inception Score actually measure
7 min readmachine-learning
FID and Inception Score get quoted like ground truth, but both inherit assumptions from ImageNet that quietly break on biomedical images. Notes on building a domain-specific feature extractor instead of trusting the Inception number.
A few years ago my brother Pavlo and I trained a GAN to synthesize cytological images — the kind a pathologist looks at to screen for cancer — ran the standard evaluation, and got a Frechet Inception Distance of 31.20. Is that good? I genuinely didn’t know, and neither did anyone reviewing the paper, because FID doesn’t have units that mean anything on their own. You only know a GAN’s FID is good relative to another FID computed the same way, on the same kind of image. We didn’t have one of those to compare against, because almost nobody runs this metric on cytological images.
That gap is what our 2023 paper ended up being about, and it’s worth unpacking properly, because the two metrics everyone quotes — FID and Inception Score — both carry an assumption that’s invisible until you point them at something that isn’t a photograph.
“Does it look good?” fails as an acceptance criterion
Eyeballing a grid of generated samples catches obvious failure — artifacts, checkerboarding, a generator that’s clearly not converged. It does not catch mode collapse reliably, because a generator that produces ten excellent, nearly identical images looks great in a grid and is nonetheless useless: it has learned to produce a narrow slice of the data distribution and nothing else. You need a number that’s sensitive to coverage, not just to whether any single sample looks plausible. That’s the whole reason FID and Inception Score (IS) exist — and the whole reason you need to know what each one is actually computing before you trust it.
What FID measures, and what it assumes
FID, from Heusel et al., 2017, doesn’t look at pixels. It runs real and generated images through a pretrained Inception-v3, pulls the activations at the final pooling layer, and reduces each image to a 2048-dimensional feature vector. Then it fits a Gaussian to the real set’s vectors and another to the generated set’s, and reports the Fréchet distance between those two Gaussians:
FID = ||m_r − m_g||² + Tr(C_r + C_g − 2(C_r C_g)^(1/2))
m and C are the mean and covariance of each feature set. Lower is better; zero means the two
Gaussians coincide.
Two things follow from that definition that get lost when people just quote the number. First, FID only sees the first two moments of the feature distribution — mean and covariance — not the whole shape. Two very different sets of images can produce matching means and covariances in feature space; FID has no way to distinguish them. Second, and more fundamental: FID is only as meaningful as the feature space it’s computed in. Inception-v3’s pool3 layer is good at describing photographs of the 1,000 ImageNet classes it was trained to classify — dogs, cars, furniture. It was never asked to represent anything about cytological or histological structure.
Inception Score inherits the same classifier, more directly
IS, from Salimans et al., 2016, is more direct about the dependency. It runs generated images through the same Inception-v3, but this time looks at the class probabilities it outputs, and combines two desiderata into one number: each individual image should produce a sharp, confident prediction (one class dominates — the image looks like something recognizable), and across the whole generated set, the predicted classes should be as varied as possible (the generator isn’t just making variations on one class). Both conditions are measured via KL divergence between the per-image conditional distribution and the marginal distribution over all generated images.
That’s a reasonable proxy — on ImageNet-like data. Point it at a cytological image and the Inception classifier has to force the image into one of 1,000 categories it has never been trained to distinguish (“tabby cat”, “sports car”), none of which bear any relation to the actual axis of variation you care about (cell morphology, nuclear staining, tissue architecture). The confidence score it produces is not meaningless noise, but it also isn’t measuring what you think it’s measuring.
Precision and recall as separate axes
A number of papers point out that collapsing “is this image realistic” and “does the generator cover the whole distribution” into one scalar hides exactly the failure mode you most need to catch. The improved precision/recall metric from Kynkäänniemi et al., 2019 splits them: it builds non-parametric manifold estimates of the real and generated feature sets using k-nearest-neighbour hyperspheres, then asks precision (“do generated samples land inside the real manifold?” — fidelity) and recall (“do real samples land inside the generated manifold?” — coverage) as two separate questions. A generator that memorizes and reproduces a narrow subset of the training set can score well on fidelity while failing coverage badly, and FID alone will not tell you that’s what happened — you’ll just see a FID that looks fine.
The domain-shift problem, with our own numbers attached
This is where the 2023 paper actually went. Instead of trusting Inception-v3’s features for the cytological task, we trained two small custom classifiers — BioCNN-1 and BioCNN-2 — on the actual image domain, and used their feature layers in place of Inception’s for both FID and IS. The result, on 64×64 cytological images: FID dropped from 31.20 (Inception features) to 0.034 (BioCNN features), IS rose from 3.52 to 3.81, and total metric computation time dropped from about two minutes to fifteen seconds, since the custom models are a fraction of Inception-v3’s size.
The honest caveat, which the headline numbers hide: that FID of 0.034 is not comparable to anyone else’s FID of anything, including our own earlier 31.20. FID and IS are only comparable when computed in the same feature space. Swap the feature extractor and you get a different metric with different units, full stop — the number getting smaller doesn’t mean our GAN improved; it means we stopped asking an ImageNet classifier to judge cytology. This is exactly the pattern the broader domain-shift literature on medical imaging describes — Stacke et al. (2020) show that CNN internal representations shift hard between natural-image pretraining and histopathology data, and the standard fix in classification work is the same one we landed on independently for evaluation: train or fine-tune on the actual domain, don’t borrow ImageNet features and assume they transfer.
The follow-up to this work, a combined metric paper from 2025, builds a single score (they call it MC) out of domain-adapted versions of IS and FID (MIS and MFID) and runs it against diffusion-synthesized histopathological images instead of GAN output — same underlying argument, newer generator family. The problem doesn’t go away when you change architectures; it’s a property of the evaluation, not the model being evaluated.
What to actually report
If you’re evaluating a generative model on anything that isn’t a natural-image photograph:
- Say what feature extractor you used. “FID: 12.4” is not a complete sentence if the model isn’t Inception-v3. If you trained a domain-specific one, say so, and don’t compare the number across papers that didn’t.
- Report more than one axis. FID and IS each compress quality and diversity into a single number; a precision/recall split, or at minimum both metrics side by side, catches failures neither one catches alone.
- Treat a domain-specific feature extractor as a real option, not a hack. It’s a few hours of training a small CNN, and it’s the difference between a metric that reflects something about your data and one that reflects how well your images coincidentally resemble ImageNet photographs.
- A single number is not a clinical claim. If the downstream use is diagnostic — same territory as the MLOps post on clinical ML — a good FID tells you the generator is a reasonable data-augmentation source. It tells you nothing about whether a classifier trained on that synthetic data is safe to point at a patient. Keep those as separate gates.
The rest of this line of work — architecture search for biomedical GANs, the combined-metric paper, the earlier dataset work — is on the research page.