Intelligent Speech Therapy Platform with Adaptive Exercises and Progress Tracking
An AI-powered speech therapy system that evaluates pronunciation accuracy, identifies phoneme-level errors, and provides personalized practice exercises to improve spoken English.
Django REST Framework • Wav2Vec2 • librosa • PyTorch • Transformers
React Router • Recharts • Lucide Icons
Google Gemini • Groq • Cerebras
- Pronunciation Assessment: Phoneme-level analysis using Wav2Vec2 embeddings and cosine similarity scoring
- ASR Validation: Whisper-based speech verification before scoring (prevents "yes man" bug)
- Multi-level Error Detection: Word, phoneme, and letter-level mistake identification with DTW alignment
- Adaptive Practice: AI-generated sentences targeting weak phonemes
- Progress Tracking: Visual dashboards showing improvement over time
- Personalized Feedback: Human-readable explanations powered by LLMs
- Reference Library: 44 English phonemes with articulation guidance
This project follows strict architectural guidelines to maintain scientific validity:
- Separation of Concerns: NLP Core handles signal processing; LLM Engine only generates feedback text
- Deterministic Scoring: All pronunciation scores use cosine similarity on embeddings
- LLM Constraints: LLMs never score pronunciation, detect phonemes, or judge speech correctness
- Precomputed Data: Reference phonemes and embeddings are cached, not regenerated per request
pronunex/
├── backend/ # Django backend
│ ├── apps/ # Django applications
│ │ ├── accounts/ # User authentication
│ │ ├── library/ # Phonemes & reference sentences
│ │ ├── practice/ # Sessions, attempts, assessment
│ │ ├── analytics/ # Progress tracking
│ │ ├── llm_engine/ # LLM feedback generation
│ │ └── sentence_engine/ # Adaptive sentence generation
│ ├── nlp_core/ # Signal processing modules
│ ├── services/ # Shared services
│ ├── config/ # Django configuration
│ └── media/ # Audio storage (gitignored)
│
├── frontend/ # React frontend
│ ├── src/
│ │ ├── components/ # Reusable UI components
│ │ ├── pages/ # Page components
│ │ ├── context/ # React context providers
│ │ ├── hooks/ # Custom React hooks
│ │ ├── api/ # API client
│ │ └── styles/ # Global styles
│ └── public/ # Static assets
│
└── docs/ # Documentation
- Python 3.10+
- Node.js 18+
- PostgreSQL (or use Supabase)
# Clone repository
git clone https://github.com/infosys-springboard-vinternship/pronunex.git
cd pronunex
# Backend setup
cd backend
python -m venv venv
venv\Scripts\activate # Windows | source venv/bin/activate (Linux/Mac)
pip install -r requirements.txt
cp .env.example .env # Configure your API keys
python manage.py migrate
python manage.py seed_data
python manage.py runserver # http://localhost:8000
# Frontend setup (new terminal)
cd frontend
npm install
cp .env.example .env
npm run dev # http://localhost:5173For detailed setup instructions, see:
- Backend README - Complete backend setup, API documentation, and architecture
- Frontend README - Frontend setup, component structure, and development guide
RESTful APIs available at http://localhost:8000/api/v1/:
/auth/- Authentication and user management/library/- Phonemes and reference sentences/practice/- Practice sessions and assessments/analytics/- Progress tracking and statistics
Admin Panel: http://localhost:8000/admin/
See Backend README for complete API documentation.
See detailed development guides:
- Backend README - Django commands, testing, deployment
- Frontend README - npm scripts, component development, styling
Please read CONTRIBUTING.md for details on our code of conduct and development process.
This project is licensed under the MIT License - see the LICENSE file for details.
Author: Abhishek Maurya
Last updated: 2026-02-07