MiniPro is a two-stage deep learning and computer vision system for authenticating Aadhaar cards and detecting forgeries using ultra-strict pixel-perfect verification.
- Stage 1 — Classification: Determines whether the uploaded document is an Aadhaar card or not (e.g. passport, driver’s license, or other ID).
- Stage 2 — Verification: If the document is classified as Aadhaar, it is compared against a reference database of genuine Aadhaar images using four strict similarity metrics.
The system also includes an optional tampering subsystem (part2/tampering_system/) with OCR, QR checking, and visual tamper detection; the main web and CLI flow use the reference-database image matcher only.
- Classification Module (Part 1): MobileNetV2 CNN for binary Aadhaar / Non-Aadhaar classification
- Verification Module (Part 2): Four-metric image matching (SSIM, ORB, Histogram, Pixel-perfect) against a reference database
- Web Interface: Flask application with drag-and-drop upload
- Reported accuracy (with test set): 95.7% overall (100% genuine detection, 90.9% forgery detection, 100% non-Aadhaar classification)
-
Stage 1: Classification (Part 1)
- Model: MobileNetV2 transfer learning (ImageNet pre-trained)
- Input: 224×224 RGB images
- Output: Binary classification (Aadhaar/Non-Aadhaar) + confidence score
- Location:
part1/model/aadhar_model.keras(20.3 MB) - Inference: ~200ms after initial model load (~8-10s first request)
-
Stage 2: Verification (Part 2)
- Method: Image matching against reference database via
part2/image_matcher.py - Metrics: 4 ultra-strict similarity measures (SSIM, ORB, Histogram, Pixel)
- Database: Genuine Aadhaar images in
part2/reference_db/(you create and populate this folder) - Decision: ALL 4 metrics must pass for verification
- Processing: ~1 second per reference image (total time depends on number of reference files)
- Method: Image matching against reference database via
- File:
analyze_document.py - Function:
analyze_document(image_path, verbose=True) - Thread-Safety: Singleton pattern with double-checked locking for model
- Returns: JSON dict with classification + verification results
- File:
app_ui.py - Framework: Flask 3.0
- Port: localhost:5000
- Routes:
GET /→ Upload pagePOST /upload→ Process image- Result page with verdict display
ORB_CONFIG = {
'nfeatures': 10000, # Maximum features
'scaleFactor': 1.05, # Fine-grained scale
'edgeThreshold': 3 # Ultra-strict edge detection
}Summary: Python 3.10+ · TensorFlow/Keras (MobileNetV2) · OpenCV · scikit-image · Flask · NumPy · (optional: pandas, matplotlib, seaborn for tests/reports)
| Category | Package | Purpose |
|---|---|---|
| Deep Learning | TensorFlow / Keras | MobileNetV2 model load and inference (Part 1) |
| Computer Vision | OpenCV (opencv-python) |
Image I/O, preprocessing, ORB features (Part 1 & 2) |
| scikit-image | SSIM (structural similarity) in Part 2 image_matcher.py |
|
| Numerical | NumPy | Array operations |
| Web | Flask | Web server and routes |
| Werkzeug | Secure filename handling for uploads | |
| Part 1 (training) | Pillow, matplotlib, scikit-learn | Data loading, viz, utilities (part1/requirements.txt) |
| Testing / reports | pandas, matplotlib, seaborn | Used by run_comprehensive_test.py for reports |
- Required: Python 3.10+ (3.13+ recommended)
- Virtual environment: Use
.venvorvenv(e.g.python -m venv .venv)
MiniPro/
│
├── analyze_document.py # Unified pipeline (Part 1 + Part 2 image matcher)
├── app_ui.py # Flask web application
├── run_comprehensive_test.py # Batch test runner (TestImages/ → reports)
├── README.md # This file
│
├── part1/ # Classification Module
│ ├── model/
│ │ └── aadhar_model.keras # MobileNetV2 trained model (20.3MB)
│ ├── dataset/ # Training/testing data (not required for execution)
│ │ └── README.md
│ ├── src/
│ │ ├── config.py # Paths, image size, training config
│ │ ├── train_model.py # Model training script
│ │ ├── predict_image.py # Standalone prediction
│ │ └── select_image.py # Image selection utilities
│ ├── requirements.txt # Part 1 dependencies (TensorFlow, OpenCV, etc.)
│ └── __init__.py
│
├── part2/ # Verification Module
│ ├── image_matcher.py # 4-metric matching engine (used by main pipeline)
│ ├── reference_db/ # Reference Aadhaar images (create and add JPG/PNG; not in repo)
│ └── tampering_system/ # Optional: OCR, QR, visual tamper detection (separate pipeline)
│ ├── app.py # Standalone tampering app
│ ├── verification_pipeline.py
│ ├── tamper_verifier.py
│ ├── ocr_extractor.py
│ ├── qr_checker.py
│ ├── quality_analyzer.py
│ ├── visual_tamper_detector.py
│ └── sample_test.py
│
├── templates/ # Web UI templates
│ ├── upload.html # File upload page
│ └── result.html # Result display page
│
├── uploads/ # Temporary file storage (gitignored)
└── .gitignore
- Windows 10/11 (PowerShell)
- Python 3.13+ installed
- 8GB RAM minimum (16GB recommended)
-
Navigate to project directory (e.g. MiniPro root):
cd c:\dev\AdhaarVerification\MiniPro
-
Create and activate a virtual environment (recommended):
python -m venv .venv .venv\Scripts\Activate.ps1 -
Install dependencies:
- Part 1 (classification) and Part 2 (image matcher) + web app:
pip install -r part1/requirements.txt pip install scikit-image flask werkzeug
Or in one go:
pip install tensorflow opencv-python numpy scikit-image flask werkzeug pillow matplotlib scikit-learn
-
Reference database: Create
part2/reference_db/and add genuine Aadhaar reference images (JPG/PNG). The verification step compares uploads against these images. -
Run web application:
python app_ui.py
-
Access in browser:
http://localhost:5000/
- Open browser → http://localhost:5000/
- Upload image → Drag & drop or click browse (JPG/JPEG/PNG, max 16MB)
- View results → Verdict + 4 metric scores + best match reference
from analyze_document import analyze_document
result = analyze_document('path/to/image.jpg', verbose=True)
# Returns:
# {
# 'is_aadhaar': bool,
# 'classification_confidence': float,
# 'match_status': 'MATCH_FOUND' | 'FORGED' | 'NOT_APPLICABLE',
# 'best_match_file': str or None,
# 'match_scores': {
# 'ssim': float,
# 'orb': int,
# 'hist': float,
# 'pixel': float
# }
# }| Condition | Verdict | UI Display |
|---|---|---|
is_aadhaar == False |
Not Aadhaar | ❌ Gray badge: "Not an Aadhaar Card" |
is_aadhaar == True AND match_status == MATCH_FOUND |
Verified | ✅ Green badge: "VERIFIED — All Strict Checks Passed" |
is_aadhaar == True AND match_status == FORGED |
Forged | 🚫 Red badge: "FORGED / TAMPERED — Failed Strict Verification" |
For Genuine Aadhaar Cards:
- Only verified if the image matches one of the reference database images (near pixel-perfect)
- Must be pixel-perfect match (99.99%+ similarity)
- Even JPEG re-compression may cause rejection
For Forged Documents:
- Any tampering detected (photo change, field edit, etc.)
- Even 1 changed pixel triggers FORGED
- Cropped/resized images rejected
For Non-Aadhaar Documents:
- Passport, driver's license, ID cards rejected
- Classification happens before verification
- Fast rejection (~200ms)
run_comprehensive_test.py: Scans theTestImages/folder (expected subfolders:Adhaar_test,Forged,Not_adhaar_test) and generates reports inreports/. Requires pandas, matplotlib, seaborn.
| Metric | Result |
|---|---|
| Genuine Accuracy | 100.0% (30/30) |
| Forgery Detection | 90.9% (40/44) |
| Non-Aadhaar Classification | 100.0% (20/20) |
| Overall Accuracy | 95.7% |
| Predicted: Genuine | Predicted: Forged | Predicted: Non-Aadhaar | |
|---|---|---|---|
| Expected: Genuine | 30 | 0 | 0 |
| Expected: Forged | 2* | 40 | 2 |
| Expected: Non-Aadhaar | 0 | 0 | 20 |
*2 forged images classified as genuine because they were identical to reference database images (duplicate issue)
def analyze_document(image_path: str, verbose: bool = True) -> dict:
"""
Unified pipeline for Aadhaar document analysis.
Args:
image_path: Path to image file
verbose: Print progress messages
Returns:
Dictionary with classification + verification results
Workflow:
1. Load & classify image (MobileNetV2)
2. If not Aadhaar → return NOT_APPLICABLE
3. If Aadhaar → run 4-metric verification
4. Return unified result
"""class AadhaarClassifier:
def predict(self, image_path: str) -> tuple:
"""
Returns: (is_aadhaar: bool, confidence: float, raw_score: float)
Model: MobileNetV2
Input: 224×224 RGB
Output: Sigmoid probability (<0.5 = Aadhaar, >0.5 = Not Aadhaar)
"""def match_with_reference(image_path: str, verbose: bool = True) -> dict:
"""
Match against all reference images in part2/reference_db/ using 4 metrics.
Returns:
{
'match_status': 'MATCH_FOUND' | 'FORGED',
'best_match_file': str or None,
'scores': {
'ssim': float (0.0-1.0),
'orb': int (0-10000+),
'hist': float (0.0-1.0),
'pixel': float (0.0-1.0)
}
}
Processing time: ~1 second per reference image (sequential).
"""| Stage | First Request | Subsequent |
|---|---|---|
| Model Loading | 8-10 seconds | 0ms (cached) |
| Classification | 200ms | 200ms |
| Verification (per reference) | ~1s | ~1s |
| Total (N refs) | Model load + N×~1s | N×~1s |
- Model in memory: ~500MB
- Peak processing: ~1GB
- Idle: ~200MB
- Parallel reference processing → reduce 70s to ~10s
- GPU acceleration → 10x faster ORB matching
- Pre-computed reference features → eliminate redundant computation
- Early stopping → return on first match
- Processing time: Scales with reference DB size (~1s per reference image, sequential)
- Reference database: You must create and populate
part2/reference_db/; larger sets improve coverage but increase runtime - Ultra-strict thresholds: High false rejection rate (~60-80% legitimate variations rejected)
- JPEG intolerance: Re-compressed images fail verification
- No OCR: Cannot extract/validate text fields (name, DOB, number)
- No QR validation: Digital signature not checked
- No batch processing: One image at a time
| Aspect | Choice | Reason |
|---|---|---|
| False Positives | Minimize (near-zero) | Security over convenience |
| Thresholds | 99.99% (pixel-perfect) | Forensic-level verification |
| Reference Database | Image-based (not OCR) | Simpler, no text extraction needed |
| Processing Speed | Slow (70s) | Accuracy over speed |
| Model Size | 20MB (MobileNetV2) | Deployment-friendly |
- Siamese Network: Learn similarity instead of hand-crafted metrics
- Perceptual Hashing: Tolerate JPEG compression (pHash/dHash)
- OCR Integration: Extract + validate text fields (name, DOB, number)
- QR Code Validation: Decode digital signature
- Parallel Processing: Multi-threaded reference comparisons
- Adaptive Thresholds: Quality-based threshold adjustment
- Expand Reference DB: 1000+ images for better coverage
- API Mode: RESTful endpoints for integration
- Batch Processing: Multiple images at once
- Real-time Video: Live camera feed analysis
- JPEG/JPG: Most common, but re-compression may cause rejection
- PNG: Lossless, best for pixel-perfect matching
- Maximum size: 16MB
- Recommended: High-quality scans (300+ DPI)
- Purpose: Ground truth for Stage 2 verification. The pipeline compares the uploaded Aadhaar image against every image in this folder.
- Setup: Create the folder
part2/reference_db/if it does not exist. Add genuine Aadhaar card images (JPG, JPEG, or PNG). The folder is not included in the repo; you supply your own reference set. - Count: Any number of images; more references improve coverage but increase processing time (~1s per reference).
- Adding new: Copy images into the folder; no config changes needed.
- Purpose: Temporary storage during processing
- Cleanup: Overwritten on subsequent uploads
- Gitignored: Not tracked in version control
| Error | Cause | Solution |
|---|---|---|
Model not found |
Missing aadhar_model.keras |
Ensure model file exists in part1/model/ |
Reference DB not found |
Missing reference_db/ |
Ensure folder exists with images |
Image load failed |
Corrupted file | Use valid JPG/PNG |
413 Error |
File too large | Reduce size below 16MB |
Memory error |
Insufficient RAM | Restart app, close other programs |
Slow processing |
Normal behavior | 70s per image is expected |
# Enable verbose output
result = analyze_document('image.jpg', verbose=True)
# Prints:
# [STEP 1/2] Classifying document type...
# [STEP 2/2] Comparing with reference database...
# [ULTRA-STRICT MATCHER] Scores: SSIM=..., ORB=..., etc.- File extension whitelist (jpg, jpeg, png)
- Size limit enforcement (16MB)
- Secure filename handling (Werkzeug
secure_filename) - No code execution from uploads
- No data persistence: Images deleted after processing
- No logging: Sensitive data not stored
- Local processing: No external API calls
- Session-based: Results cleared on new upload
- Enable HTTPS (SSL/TLS)
- Add authentication (Flask-Login)
- Implement rate limiting (Flask-Limiter)
- Use CSRF tokens
- Deploy behind reverse proxy (Nginx)
- Use production WSGI server (Gunicorn)
- Add audit logging
- Implement backup/recovery
# Already configured
python app_ui.py
# Access: http://localhost:5000/# Install Gunicorn
pip install gunicorn
# Run with 4 workers
gunicorn -w 4 -b 0.0.0.0:5000 app_ui:app
# With Nginx reverse proxy
# /etc/nginx/sites-available/aadhaar
server {
listen 80;
server_name your-domain.com;
location / {
proxy_pass http://127.0.0.1:5000;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}FROM python:3.13-slim
WORKDIR /app
COPY . .
RUN pip install -r part1/requirements.txt
EXPOSE 5000
CMD ["python", "app_ui.py"]Problem: "Model not found" error
Solution: Verify part1/model/aadhar_model.keras exists (20.3 MB)
Problem: "Could not load model" error
Solution: Install TensorFlow 2.20.0: pip install tensorflow==2.20.0
Problem: Takes >2 minutes per image
Solution: Normal behavior (one comparison per reference image). Consider parallel processing for speed.
Problem: All images marked as FORGED
Solution: Expected with ultra-strict mode. Only pixel-perfect matches pass. Check if test images exist in reference database.
Problem: Cannot access localhost:5000
Solution: Check if Flask is running. Ensure no firewall blocking port 5000.
Problem: File upload fails
Solution: Check file size (<16MB) and format (JPG/PNG only).
- MobileNetV2: Sandler et al., "MobileNetV2: Inverted Residuals and Linear Bottlenecks" (2018)
- Transfer Learning: ImageNet pre-trained weights
- TensorFlow/Keras: Google Brain Team
- OpenCV: Intel, Willow Garage, Itseez
- scikit-image: scikit-image development team
- Flask: Pallets Projects
- Training dataset: Custom Aadhaar card images (proprietary)
- Reference database: 71 genuine Aadhaar samples
Proprietary - For educational/research purposes only.
Repository: Aadhar-Part1 (GitHub)
Owner: Likhithagowda25
Branch: main
python app_ui.pyfrom analyze_document import analyze_document
result = analyze_document('test.jpg', verbose=True)
print(result['match_status'])- SSIM: ≥ 0.9999
- ORB: ≥ 500
- Histogram: ≥ 0.9999
- Pixel: ≥ 0.9999
- First request: ~80 seconds
- Subsequent: ~70 seconds
- Python 3.13+
- 8GB RAM minimum
- Windows 10/11
Last Updated: February 2026
Version: 2.0 (Ultra-Strict Pixel-Perfect Mode)
Project: MiniPro (Aadhaar Verification)