Home/About

AI Architect & Engineering Manager · 10 years

Shivank Shukla

I build production AI systems and the infrastructure underneath them — and I publish what breaks, because that's the part nobody writes about.

Book a consultation call

What I've built

Ten years in distributed systems and cloud infrastructure, currently owning end-to-end architecture for a multi-tenant GenAI platform on AWS. Rust, Go, Python and Java across LLM serving, RAG, agentic workflows and MLOps.

Retrieval

12M documents, measured

RAG on Amazon OpenSearch Serverless and Bedrock Knowledge Bases, with a Rust ingestion path chunking and embedding at 8M records/hour, plus an evaluation harness scoring retrieval precision, groundedness and hallucination rate against a curated golden set.

12M+ documents · 8M records/hour ingestion

Self-hosted

40M vectors on Kubernetes

Qdrant with HNSW tuning and payload filtering, hybrid retrieval combining BM25 and dense similarity with cross-encoder reranking, open-weight models served via vLLM with continuous batching, evaluation gates wired into CI.

p95 under 25ms at 1.5K QPS · 3.2× serving throughput

Gateway

A model gateway in Rust

Centralised LLM gateway fronting Bedrock — Claude, Llama, Titan — with model routing, token streaming, per-tenant rate limiting, request tracing and cross-region failover. Consolidated three fragmented team integrations into one.

40K+ concurrent connections · sub-100ms overhead

Cost

Spend down while volume grew

Token-level cost attribution per team and use case via CloudWatch and Cost Explorer tagging, prompt caching and model right-sizing. Separately, moved high-volume batch inference to EKS with vLLM and Karpenter across G6 and Inferentia2.

$180K annual spend removed · 55% lower per-token cost on batch

Agents

Agents with approval gates

Multi-step agent systems on Bedrock Agents and Step Functions: tool-calling contracts, MCP server integrations to internal APIs, human-in-the-loop gates on high-risk actions, and replay and audit trails on every decision path.

In production

Platform

Terraform for the whole footprint

Bedrock access policies, EKS clusters, OpenSearch collections, VPC endpoints, IAM boundaries — with ArgoCD GitOps delivery and observability through Prometheus, Grafana and OpenTelemetry. Bedrock Guardrails, SageMaker Clarify, CloudTrail audit logging.

SOC 2 aligned · all model traffic on PrivateLink

Before this

Engineering Manager at a European telecommunications group — led a team of engineers across platform and applied AI, set engineering standards, ran architecture review, and drove adoption across five business units. Built a Rust and Apache Arrow pipeline processing 10M+ records/minute, and a production HTTPS server handling 50K+ concurrent connections at sub-millisecond response.

Technical Lead at a cloud consultancy — migrated AWS CloudFormation to Terraform HCL across 10 environments, and rewrote the CloudFormation automation tooling in Rust: 3× faster execution, and it eliminated the memory bugs behind 15% of failed deployments.

Credentials

  • HashiCorp Terraform Associate
  • AWS Certified Solutions Architect – Associate (in progress)
  • AWS Certified AI Practitioner (in progress)
  • PostgreSQL Certification
  • DevOps Proficiency & Expertise, Level 3 — Triplebyte
  • CodeWars 4 kyu — top 10% globally, data structures and algorithms
  • B.Tech, Electronics & Communication

Writing

I publish the architecture and the failures at medium.com/@aiinfra — agentic RAG in Postgres with the raw routing payloads, why 95%-accurate agents finish 100-step jobs 0.6% of the time, and building a semantic cache in Rust with the one metric that tells you it's returning wrong answers.

Code at GitHub. Shorter notes at @c199benzene.

Working together starts with a call. Thirty minutes, no charge — you describe what's breaking and I tell you whether I can help. If an audit is the right next step, it's five days at a fixed price.

Book a consultation call

Last updated 13 September 2026