Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
Abstract
The growing realism and accessibility of manipulated and generated faces threaten the trustworthiness of digital media. To detect such forgeries, deepfake detectors based on vision foundation models have shown promising performance, but they typically rely on a single pretrained representation and are prone to overfitting to particular training distributions. To improve generalization to unseen forgeries, we propose UCF-Net, an uncertainty-aware cascaded fusion network that harnesses CLIP's language-aligned semantic priors and DINO's self-supervised visual-structure priors. UCF-Net extracts hierarchical features across Transformer depths, uses layer-wise expert aggregation to adaptively combine each encoder's multi-level cues, and performs weighted fusion of the resulting representations based on entropy-derived uncertainty. We further consolidate public deepfake datasets into a unified benchmark of approximately 4M images and construct a separate cross-generator evaluation set with over 8K face images from eight recent generators. On the unified benchmark, UCF-Net achieves the best mean AUC among the evaluated methods in both in-domain and cross-domain evaluations. On the cross-generator set, it adapts effectively with limited target-domain data, although zero-shot transfer remains challenging.
Motivation
CLIP and DINO capture distinct evidence patterns
Their distinct evidence patterns and failure modes motivate uncertainty-aware fusion.
Method
Harnessing CLIP and DINO priors with uncertainty-aware cascaded fusion
UCF-Net preserves hierarchical CLIP and DINO cues, adaptively aggregates multi-level features, and fuses the resulting representations according to uncertainty.
Retain hierarchical representations
Keep multi-level CLIP and DINO features.
Adaptively aggregate multi-level features
LEA adaptively aggregates multi-level features.
Balance CLIP and DINO evidence
UAF balances the two representations by feature uncertainty.
UCF-Net
Benchmark
A unified benchmark for evaluating generalization in deepfake detection
The unified benchmark and a separately constructed cross-generator image evaluation set support in-domain, cross-domain, and cross-generator evaluation.
Overview of the unified deepfake benchmark and the separately constructed cross-generator evaluation set
A unified deepfake benchmark
The unified benchmark covers four forgery categories under binary real/fake supervision: face swapping (FS), face reenactment (FR), entire face synthesis (EFS), and face editing (FE).




Representative examples of the four forgery categories in Figure 3(a).
Eight recent image generators
A separate set for zero-shot transfer and few-shot adaptation.








Dataset construction
From provenance-bearing records to audited face crops
- Data acquisitionVerify provenance and remove duplicates.
- Face processingDetect, align, crop, and re-detect faces.
- Human auditAudit validity, occlusion, realism, and clarity.
Main results
Generalization beyond the training distribution
UCF-Net is shown with a solid reference curve, while compared methods use dashed curves. Click a legend item to show or hide a method, and hover or focus any point to read its exact AUC.
Best mean AUC across six in-domain datasets.
In-domain mAUC
Click a legend item to show or hide a method; hover or focus a point for exact AUC.
Scale: 0–100 AUC (%)
Method ranking
Best mean AUC across held-out datasets and DF40 forgery categories.
Cross-domain AUC
Select a method in the legend to show or hide it. Focus a point for its exact AUC.
Scale: 0–100 AUC (%)
Method ranking
5 fake + 5 real images per generator; evaluation uses the remaining generated images and 8,807 real images sampled in a dataset-balanced way from the original test split.
Cross-generator AUC
Click a legend item to show or hide a method; hover or focus a point for exact AUC.
Scale: 0–100 AUC (%)
| Method | CDF | DFFD | DFDCP | FF++ | DF40 | MFFI | mAUC |
|---|---|---|---|---|---|---|---|
| Xception | 99.26 | 98.48 | 90.93 | 76.03 | 83.89 | 78.94 | 87.92 |
| RECCE | 99.38 | 99.69 | 89.04 | 78.70 | 88.60 | 75.29 | 88.45 |
| SPSL | 98.97 | 98.87 | 93.49 | 82.32 | 81.92 | 81.43 | 89.50 |
| CLIP | 94.83 | 99.39 | 86.78 | 74.61 | 87.45 | 82.60 | 87.61 |
| Effort | 99.52 | 99.95 | 96.51 | 85.14 | 92.56 | 85.60 | 93.21 |
| DINOv2 | 80.28 | 98.11 | 83.50 | 66.08 | 83.75 | 71.39 | 80.52 |
| DFF-Adapter | 99.85 | 99.96 | 97.31 | 92.10 | 91.60 | 88.31 | 94.86 |
| UCF-Net | 99.74 | 99.99 | 97.88 | 90.55 | 92.90 | 90.92 | 95.33 |
| Method | UADFV | DFF | DFDC | DF40-FS | DF40-FR | DF40-EFS | DF40-FE | mAUC |
|---|---|---|---|---|---|---|---|---|
| Xception | 91.37 | 53.14 | 75.16 | 66.00 | 89.90 | 45.96 | 46.91 | 66.92 |
| RECCE | 85.59 | 72.62 | 73.37 | 76.27 | 64.50 | 42.89 | 48.04 | 66.18 |
| SPSL | 94.29 | 56.72 | 73.02 | 60.65 | 74.26 | 61.89 | 49.72 | 67.22 |
| CLIP | 95.46 | 85.52 | 77.06 | 95.21 | 89.04 | 74.65 | 93.62 | 87.22 |
| Effort | 98.24 | 88.91 | 88.61 | 95.51 | 88.39 | 69.04 | 93.67 | 88.91 |
| DINOv2 | 85.67 | 58.12 | 72.71 | 94.35 | 74.62 | 61.51 | 95.04 | 77.43 |
| DFF-Adapter | 99.08 | 81.97 | 87.63 | 92.45 | 90.11 | 77.78 | 95.37 | 89.20 |
| UCF-Net | 97.49 | 89.72 | 90.01 | 97.85 | 93.38 | 78.10 | 98.52 | 92.15 |
| Method | 0-shot | 5-shot | 50-shot | 100-shot |
|---|---|---|---|---|
| Xception | 46.72 | 77.41 | 91.85 | 95.66 |
| RECCE | 32.80 | 77.06 | 94.08 | 91.45 |
| SPSL | 60.02 | 82.72 | 90.96 | 96.11 |
| Effort | 37.23 | 83.03 | 97.10 | 98.49 |
| DFF-Adapter | 43.21 | 89.23 | 97.07 | 97.27 |
| UCF-Net | 40.90 | 91.36 | 98.24 | 98.81 |
Ablation study
What contributes to the gain?
Each card changes one design choice while the remaining training setup stays fixed. The bars report cross-domain mAUC.
CLIP–DINO encoder pair
The heterogeneous CLIP–DINO pair performs best.
Layer-wise aggregation
LEA adaptively aggregates multi-level features across Transformer depths.
Fusion strategy
UAF outperforms fixed and parameter-heavy alternatives.
LoRA rank
Increasing adaptation capacity does not always improve transfer.
Training data scale
Scale of Training Data
Cross-domain mAUC under 10K, 1M, and 2M training settings. The 10K and 1M sets are proportional samples from the full training split; 1M is half of the full split, while 2M denotes the complete approximately 2.2M-image training split.
Cross-domain mAUC (%). Click a legend item to show or hide a method; hover or focus a point for the exact value.
View exact training-scale values
| Method | 10K | 1M | 2M |
|---|---|---|---|
| Xception | 69.81 | 68.14 | 66.92 |
| RECCE | 63.18 | 63.01 | 66.18 |
| SPSL | 65.88 | 70.34 | 67.22 |
| Effort | 83.37 | 88.79 | 88.91 |
| DFF-Adapter | 85.12 | 89.70 | 89.20 |
| UCF-Net | 85.58 | 91.46 | 92.15 |
Qualitative evidence
Seeing what the model sees
Grad-CAM and feature-space visualizations provide qualitative evidence for how CLIP and DINO representations are fused.
More concentrated facial responses in the shown examples
On the shown examples, UAF produces more concentrated facial responses than the compared fusion strategies.
More separated clusters in the plotted representations
The 2-D t-SNE view shows more separated real/fake clusters in the plotted representations; this is not detection AUC.
Linear separability in the plotted 2-D embeddings
UCF-Net reaches 87.0% joint-view and 92.5% DF40-Test 2-D t-SNE linear separability; these values are not detection AUC.
CLIP and DINO attention in additional examples
Across the shown samples, CLIP and DINO exhibit different attention patterns.
Release materials
Resources & Citation
Paper, source code, pretrained models, datasets, and the project citation entry.