Local AI · Python 2025 Mockup / study

Traductor de llamadas

Live call translation, fully local. STT → LLM → voice-cloning pipeline with sub-500ms latency on a 6 GB GPU.

Why it was built

An end-to-end real-time call translation pipeline running entirely on a local RTX 4050 6 GB GPU — zero cloud inference cost, zero data leaving the machine.

The flow is Whisper (Speech-to-Text) → Ollama with Qwen2.5 (translation) → XTTS with the original speaker's cloned voice (Text-to-Speech), routed into the call via VB-Cable. End-to-end latency under 500 ms.

How it works
  • Whisper for live transcription
  • Ollama + Qwen2.5 for local translation
  • XTTS with speaker voice cloning
  • Audio routing with VB-Cable
  • Sub-500ms end-to-end latency
  • GPU RTX 4050 6 GB — 100% local
Stack
Python Whisper Ollama Qwen2.5 XTTS VB-Cable
Notes

Not a web project, but one of the most ambitious local-AI tools in the portfolio. Open source.