Krushnkumar Kathrecha Book a call

AI / ML engineer, Hyderabad

The demo worked. Production is a different system.

Most LLM projects stall in the same place. Retrieval that looked fine on a sample folder falls apart on ten years of real documents. Agents loop, or quietly fail without anyone noticing. Latency and token spend go somewhere nobody budgeted for.

That gap is what I do. Thirteen years building and running software, three of them as CTO, now spent entirely on getting language-model systems past the prototype and keeping them alive once they are there.

What I’m hired to do

Four things, and I turn down work outside them. Depth in a narrow band is the reason to hire an independent rather than an agency.

Retrieval that survives real data

RAG systems that hold their accuracy when you point them at the whole archive instead of the sample. Chunking and parsing for messy source documents, hybrid and graph retrieval, reranking, grounding checks, and an evaluation harness so you can tell whether a change made things better or worse.

Google Agent Development Kit (ADK)   LangChain / LangGraph   Haystack   Qdrant   Graph RAG   Bedrock

Agents that terminate and recover

Multi-step agentic systems built to fail safely: bounded loops, tool contracts that hold, retries and fallbacks, human handoff where it belongs, and traces detailed enough to reconstruct any run after the fact. Built with MCP and A2A where they earn their place.

LangGraph   MCP / A2A   Claude & OpenAI agent SDKs   tool design

Real-time voice and telephony AI

Conversational pipelines that run under a second end to end, which is a different engineering problem from a chat box. Streaming, barge-in, model and provider routing, latency budgeting across every hop, and graceful behaviour when one hop is slow.

vLLM   LiteLLM   streaming pipelines   telephony integration

Diagnostic on a stalled build

Two weeks, fixed fee, no commitment beyond it. I read the code and the traces, talk to the team, and hand back a written assessment: what is actually wrong, what it costs to fix, what to stop doing, and whether the thing is worth continuing at all. Some clients stop there. That is a fine outcome.

Often the fastest way to find out if we should work together

Recent work

Fraud detection system

For bank transactions to detect fraud in real time. The system uses a combination of machine learning models and rule-based algorithms to flag suspicious transactions, and it integrates with the bank's existing infrastructure for seamless operation.

OCR

Optical Character Recognition (OCR) system to extract text from scanned documents and images. The system is designed to handle various fonts and layouts, providing accurate text extraction for further processing and analysis.

Chatbot for finance system

AI-powered chatbot for financial institutions to assist customers with account-related queries, providing instant support and reducing the burden on human agents.

Quant finance

A multi-agent system for quantitative finance to analyze stock market data and predict buy, hold, or sell signals.

Live multilingual translation

Ekip IT Solutions

2025–2026

Architected an LLM-powered translation system for live conversation, where anything above a second of delay makes the product unusable. The work was mostly latency budgeting: choosing what could be streamed, what could be predicted ahead, and what had to be cut from the path entirely.

Sub-second end-to-end latency

Telephony agents in production

Ekip IT Solutions

2025–2026

Designed and shipped voice agents handling real customer calls with natural language understanding and context carried across the conversation. Alongside them, agentic systems that run multi-step workflows on their own, make decisions mid-call, and adapt when the conversation goes somewhere the script did not anticipate.

Live customer traffic

CTO, two SaaS products

BAWES · Kuwait, remote

2016–2025

Nine years with one group, the last three as CTO, owning engineering for Plugn (eCommerce SaaS) and Studenthub (student job portal). Set technology strategy, hired and ran the team, moved the platform onto microservices so AI could be introduced without rewriting everything, and kept a public SaaS standing through DDoS, spam and brute-force attacks.

10× web application performance

Achieved through React/Vite upgrades, vendor chunking, selective module loading, service workers, Cloudinary image handling, Algolia search and CloudFront edge caching — with AWS auto-scaling and Go for CPU-bound work on the backend.

How engagements run

Scope and price agreed before anything starts. No open-ended hourly billing.

Diagnostic

2 weeks · fixed fee

A written assessment of an existing AI system: failure modes, cost drivers, what to fix first, and an honest read on whether to continue. Ends cleanly with no obligation to go further.

Build

6–12 weeks · per project

I build the system and hand it over working, documented, instrumented and with evals your team can run without me. Handover is part of the scope, not an afterthought.

Fractional AI lead

2–3 days/week · monthly

Ongoing technical ownership for teams building AI without a senior person who has shipped it before. Architecture, code review, hiring input, and saying no to the expensive ideas.

Working with European teams

I have worked remotely for companies outside India since 2016, including three years running engineering for a Kuwait-based group. The practical friction of cross-border work is already solved on my side.

Five to six hours of daily overlap
IST runs ahead of CET, so my afternoon is your morning and midday. Standups, pairing and live debugging all work. This is not follow-the-sun handover.
GDPR paperwork ready
Data Processing Agreement and Standard Contractual Clauses for transfer outside the EEA, prepared in advance. Happy to work inside your infrastructure and never hold personal data at all.
Clean invoicing in EUR and GBP
Registered Indian service exporter. B2B reverse charge, so no EU VAT on my invoices. Bank transfer or SEPA.
Documentation and audit trails
Systems built with the logging, evaluation records and technical documentation you will need under the AI Act, produced as the work happens rather than reconstructed later.

Stack

AI systems

  • Google Agent Development Kit (ADK)
  • LangChain / LangGraph
  • RAG / Graph RAG
  • MCP / A2A
  • Claude & OpenAI agent SDKs
  • Haystack
  • Vector search / Qdrant
  • vLLM / LiteLLM
  • Fine-tuning / PyTorch
  • Hugging Face
  • AWS Bedrock

Languages

  • Python
  • Go
  • TypeScript
  • Rust

Data

  • PostgreSQL
  • MongoDB
  • Redis
  • Qdrant
  • Algolia

Infrastructure

  • AWS / Azure / GCP
  • Docker
  • Kubernetes
  • Terraform
  • Nginx
Google Cloud — I'm a certified Generative AI Leader

Google Cloud’s Generative AI Leader certification covers the part clients usually care about most: deciding where generative AI belongs in a business, what it will cost to run, and where the governance and risk sit. It pairs with the engineering rather than repeating it.

Verify this credential

Credentials

Every badge below links to the issuer for verification.

Also completed: Hugging Face — Fine-tuning Language Models, LLM Fundamentals, and Fundamentals of MCP. Bachelor of Engineering in Computer Engineering, Government Engineering College, Rajkot.

Tell me what stopped working

Thirty minutes, no pitch. Describe the system and where it is stuck, and I will tell you what I think is happening and whether I am the right person for it. If I am not, I will say so.