Local AI · Python
2025
Mockup / study
Traductor de llamadas
Live call translation, fully local. STT → LLM → voice-cloning pipeline with sub-500ms latency on a 6 GB GPU.
Why it was built
An end-to-end real-time call translation pipeline running entirely on a local RTX 4050 6 GB GPU — zero cloud inference cost, zero data leaving the machine.
The flow is Whisper (Speech-to-Text) → Ollama with Qwen2.5 (translation) → XTTS with the original speaker's cloned voice (Text-to-Speech), routed into the call via VB-Cable. End-to-end latency under 500 ms.
How it works
- Whisper for live transcription
- Ollama + Qwen2.5 for local translation
- XTTS with speaker voice cloning
- Audio routing with VB-Cable
- Sub-500ms end-to-end latency
- GPU RTX 4050 6 GB — 100% local
Stack
Python
Whisper
Ollama
Qwen2.5
XTTS
VB-Cable
Notes
Not a web project, but one of the most ambitious local-AI tools in the portfolio. Open source.