OmniMind: Enterprise RAG & Distributed LLM Inference Engine
A high-throughput Retrieval-Augmented Generation (RAG) platform and autonomous agent system designed for enterprise-scale semantic document synthesis and low-latency question answering.
- Engineered a hybrid vector retrieval pipeline combining dense embeddings (Hugging Face / OpenAI) with BM25 sparse keyword ranking, cutting query hallucinations by 64%.
- Built an asynchronous FastAPI microservice backend with Redis semantic caching, achieving sub-45ms p95 response times across 20,000+ daily conversational queries.
- Orchestrated containerized model deployment on AWS with Docker and Kubernetes, incorporating automated token-rate tracking and fallback LLM routing.