I break things, experiment, rewrite, and push engineering boundaries.
Backend engineer who owns AI infrastructure end-to-end. I build systems that scale, systems that stay up, and systems where the engineering decisions actually matter. No handoffs, no ticketsβif there's a problem, I fix it.
Currently at Remotus, building the entire AI backend alone: APIs, LLM pipelines, vector search, async infrastructure, you name it.
Multi-CSV ingestion β GraphQL schema β LLM-powered analytics
- Ingested unstructured data (CSV, PDF, sheets) and mapped to queryable GraphQL schema
- Built RAG pipeline achieving 98% query accuracy on structured data
- Auto-generated dashboards with custom insights and complex calculations
- Solved the hard part: normalizing messy real-world data so LLMs can reason over it
Tech: FastAPI, PostgreSQL, Qdrant, LangChain, GraphQL, AWS SageMaker
Because off-the-shelf solutions couldn't keep up
- Dropped response times from 5s β 2s through parallel streaming
- Built on async WebSocket architecture, handling high concurrency
- Real-time token streaming for better UX
Tech: FastAPI, WebSockets, async Python, event-driven architecture
Live trading signal generation using Fyers API
- Real-time options chain processing every 60 seconds
- Computed OI signals, PCR, IV skew, max pain, gamma blast detection
- Dashboard + React-based strike picker for options trading decisions
- Built because existing tools didn't give me the signal I needed
Tech: FastAPI, Fyers API, Redis, React, real-time data pipelines
Handling 10,000+ concurrent users on a single web-space
- Architected CRDT-based design using YJS for race condition elimination
- Solved data sync issues in high-concurrency environments
- Maintained 99.9% uptime under production load
Tech: WebSockets, CRDT, async Python, connection pooling
- Automated model training and hosting workflow
- 30% cost reduction through intelligent resource management
- End-to-end pipeline from data ingestion to deployment
Tech: AWS SageMaker, Python, MLflow, Unsloth
- Scaled Qdrant to 8M vectors maintaining 90%+ retrieval accuracy
- 50% reduction in storage footprint
- 30% faster semantic search through index tuning and chunking strategy
Tech: Qdrant, embedding optimization, hybrid search strategies
Languages: Python, JavaScript, Java, C++, Bash
Backend: FastAPI, Django, Node.js, async I/O, microservices
Databases: PostgreSQL, MongoDB, Neo4j, Qdrant, Redis, ClickHouse
AI/ML: LLMs (OpenAI, Anthropic, open-source), RAG, prompt engineering, function-calling, LangChain, vector databases, model serving
Infrastructure: AWS (EC2, SageMaker, S3), Docker, Kubernetes, CI/CD, Terraform
Distributed Systems: Event-driven architecture, async queues (BullMQ, Celery), WebSocket streaming, load balancing, connection pooling
Tools & Practices: Git, GitHub, system design, SOLID principles, API design (REST, GraphQL), observability, monitoring
- Agentic AI architectures β multi-step orchestration, tool use, state management
- LLMOps & evals β building evaluation frameworks for production LLM systems
- Advanced prompt engineering β reasoning, chain-of-thought, structured outputs
- Production AI systems β reliability, observability, cost optimization at scale
- Built a hierarchical agent workflow where parent tasks split into specialized sub-agents that further decompose work based on input
- Solved infinite context memory management in LLM pipelines (one of the hard problems nobody talks about)
- Optimized LLM API costs by 40% through intelligent caching, prompt compression, and request deduplication
- Designed event-driven streaming architectures with Kafka and Redis Streams for real-time inference
Actively contributing to projects that solve real problems:
- Cheshire Cat β Refactored backend endpoints, improved architecture
- db-agent β LLM integration and hallucination reduction (8+ PRs)
- db-pilot β LLM integration and model serving improvements
- GitHub Issue Metrics β Documentation refactoring
- PyVista β Fixed execution path issues
- SkriptLang β Updated build scripts and core infrastructure
- AcadVault β Backend improvements
Master's in Information Technology
Dhirubhai Ambani Institute of Information and Communication Technology (DAIICT)
CPI: 8.0/10.0 (Jul 2022 - Jun 2024)
Bachelor's in Computer Application
The Maharaja Sayajirao University of Baroda
CPI: 8.53/10.0 (Jul 2019 - May 2022)
- Systems, not features β Architecture decisions matter more than code velocity
- Ownership mindset β I debug at 2am because the problem exists, not because someone asked
- Production first β Code that doesn't run in production isn't code
- Measure before optimizing β I use metrics, logs, and profiling to find real bottlenecks
- Trade-offs over perfection β Fast and working beats slow and perfect
- Learn from others' code β Open source contributions teach faster than writing solo
β
Systems I build stay up without constant babysitting
β
Architecture decisions that compound in value over months
β
Teammates saying "this code is easy to understand"
β
Finding the bottleneck in 30 minutes instead of 3 hours
β
Building something that didn't exist before
Backend Engineer | AI Infrastructure | LLM Integration roles at:
- Well-funded startups building AI systems (not just wrapping APIs)
- Product companies solving real problems at scale
- Organizations where engineering decisions matter
Ideal setup: Remote, fast-moving team, ownership from day one, hard problems that compound learning
Have an interesting problem? Building something ambitious? Hit me up.
- Email: mihirkohli2001@gmail.com
- LinkedIn: linkedin.com/in/mihirkohli
- GitHub: github.com/MihirKohli
- Phone: +91-9737635737
"I break things, experiment, rewrite, and push boundaries. That's just how I work."


