# Naman Kundra > Full-stack and AI systems engineer. Agents you can trust in production, distributed systems that survive failure, and the products around them. ECE @ Thapar Institute. Open to internships and full-time roles. Email: naman@naman.sbs ## Links - GitHub: https://github.com/naman777 - LinkedIn: https://www.linkedin.com/in/naman-kundra-850209281/ - Resume (PDF): /resume.pdf on this site ## Flagship projects ### Nightshift: An AI on-call engineer that survives its own crashes. - Area: Multi-agent AI · Reliability - Why: When an alert fires at 3 a.m., the first hour goes into correlating metrics, logs and deploys by hand. I wanted an investigator that does that legwork, backs every claim with evidence, and can never take a risky action without a human. - How: A commander agent plans hypotheses and hands one precise question to each of four specialists (metrics, logs, changes, code). They work in parallel and write to an evidence board. Risky fixes go through a trust gate. The whole investigation is a Temporal workflow, so killing the worker mid-incident just resumes it, with no repeated LLM calls. It was also onboarded onto real Linux hosts, including an unmodified WordPress stack, with auto-discovered service graphs. - Results: 96% root-cause accuracy on held-out incidents (real LLM); 0% unsafe actions across every run; 100% claims grounded in evidence; 50 injected production failures in the public benchmark - Stack: Python, Temporal, Multi-agent LLMs, Next.js, Loki, Linux - Live demo: https://nightshift.naman.sbs - Source: https://github.com/naman777/nightshift ### LogLens: A log-intelligence model, trained from scratch, that runs on a laptop CPU. - Area: ML from scratch · Observability - Why: An incident produces thousands of log lines, and most of them are noise. Sending all of it to an LLM is slow and expensive. LogLens flags the incident window and returns just the handful of lines that explain it, with zero LLM API calls. - How: Logs are parsed, masked (IPs, paths, durations) and BPE-tokenized. A 33.6M-parameter line encoder pretrained on 14 log systems feeds a window model with anomaly and per-line suspicion heads. It is exported to int8 ONNX with an embedding cache, and exposed as a CLI, an HTTP API and an agent tool. - Results: 140× fewer tokens for an on-call agent to read; 0.94 root-cause Recall@5 (vs 0.17 random); 0.954 BGL anomaly F1 (vs 0.545 baseline); 37.5 MB int8 ONNX model, CPU-only - Caveat: Root-cause numbers are on synthetic lab incidents. The repo documents where it fails, including zero-shot transfer. - Stack: PyTorch, Transformers, ONNX Runtime, FastAPI, Python - Source: https://github.com/naman777/loglens ### Fathom: Chat with your documents, and see exactly where every answer came from. - Area: RAG · Agents · Evaluation - Why: Most "chat with your PDF" demos can't tell you where an answer came from, or whether retrieval works at all. Fathom cites the exact passages, streams its reasoning trace, and ships with an eval harness, so quality is measured, not assumed. - How: Hybrid retrieval (Postgres full-text + pgvector, fused with RRF) is followed by an LLM reranker. An agent loop splits multi-part questions, searches in parallel, checks whether the evidence is enough, and follows up once if it isn't. Answers stream over SSE with citation chips. It has per-IP rate limits and an injection scan on ingest. - Results: 0.99 answer accuracy on real docs (IETF RFCs); 4× multi-part full-hit with the agent loop (0.16 → 0.64); 72 golden questions in the eval harness; $0.0009 average cost per question, live - Stack: FastAPI, Postgres + pgvector, Next.js, OpenAI, SSE - Live demo: https://fathom.naman.sbs - Source: https://github.com/naman777/Fathom ### Foreman: A fault-tolerant distributed job scheduler you can try live. - Area: Distributed systems - Why: I wanted to understand what really happens inside a scheduler: leasing work, heartbeats, retries, timeouts, and what to do when a worker dies halfway through a job. So I built one and put it on the public internet. - How: A TypeScript coordinator schedules jobs with Redis locks and tracks workers by heartbeat. Workers run each job in its own Docker container, upload /output to S3-compatible storage, and stream status over WebSocket to a Next.js dashboard. Visitors can run guided jobs without an account; the public API is locked to fixed scenarios. - Results: 3 Docker workers in the Compose stack; 5 guided failure scenarios, public; 9 failure paths covered by the smoke suite; 0 sign-ups needed to try it - Stack: TypeScript, Node.js, PostgreSQL, Redis, Docker, MinIO, WebSockets - Live demo: https://foreman.naman.sbs - Source: https://github.com/naman777/Foreman ### C++ Load Balancer: A TCP/HTTP load balancer in C++17, with no external libraries. - Area: Systems · Networking - Why: To learn networking below the framework layer: raw sockets, thread pools, and the real trade-offs between balancing algorithms. It later became the first real host Nightshift was tested against. - How: An accept loop hands sockets to a worker pool. Each connection reserves a backend under one mutex using least-connections, round-robin, IP-hash or Rendezvous hashing, then proxies bidirectionally with select(). A health checker probes backends every 5 s, a stats thread serves JSON, and SIGHUP reloads config live. - Results: 75K requests/s sustained (one backend alone: ~40–45K); +50 µs added p50 latency over a direct connection; 50 ms to eject a SIGKILLed backend, 0–1 failed requests; 25.06% keys remapped when adding a backend (ideal 25%) - Caveat: Measured on WSL2 loopback with the load generator on the same machine. It holds 950 concurrent connections and crashes near 1,000 (select() fd limit); the repo documents the epoll fix. - Stack: C++17, POSIX sockets, Threads, HTTP, Next.js dashboard - Live demo: https://load-balancer-cpp.vercel.app - Source: https://github.com/naman777/Load-Balancer-CPP ## Other projects - [Operator](https://github.com/naman777/operator): A personal work-execution agent. It analyses job postings against your own resume evidence, drafts cited applications, and waits for human approval before any action. Result: 40-case regression suite, including prompt-injection cases. Stack: Python, Temporal, MCP, Next.js. - [HirePrism](https://hireprism-naman.streamlit.app): Placement intelligence over 654 real campus offers: a DuckDB/Parquet pipeline, anomaly detectors, and a LangGraph agent that answers questions in plain English. Result: 305 tests · 93% coverage · 8 data-backed findings. Stack: DuckDB, LangGraph, Streamlit. - [LawVista](https://lawvista.naman.sbs): Legal research platform built for the Smart India Hackathon finals: RAG legal chat, document analysis, statute identification and case-outcome prediction. Result: Built for SIH 2024 finals (top 2.7% nationally). Stack: React, Node.js, Django, Pinecone. - [Examina](https://exam-portal-next-js-eight.vercel.app): The exam platform used for ACM Thapar recruitment. Redis caching and efficient backend communication keep it fast under heavy concurrent load. Result: 10,000+ concurrent users · sub-200 ms responses. Stack: Next.js, PostgreSQL, Redis. ## Experience ### Full Stack Developer, Oracia.ai (E-VNTS), Mar 2025 – Jul 2026 Remote · Delaware, USA - Delivered scalable features for an AI real-estate platform, increasing conversions by 25%. - Optimized RESTful APIs with Node.js and PostgreSQL, reducing latency by 50%. - Implemented AI-driven automation with LangGraph and LangChain, cutting manual processes by 40%. - Stack: Next.js, Node.js, PostgreSQL, LangGraph, LangChain What I worked on at Oracia, as a product walkthrough: 1. Command center: A realtor's day starts here: every stage of the funnel, from entry consent and qualification to offers and closing, with the leads that need a human flagged up front. 2. Lead pipeline: Every lead in one table: deal stage, deal amount and an AI-estimated chance of closing, with the next step one click away. The sales panel puts intent, budget and urgency right next to the live chat. 3. Conversation intelligence: While an agent chats with a client, Oracia reads the conversation: buyer emotion, sale stage and deal cycle, plus upcoming and past steps extracted automatically. Ask it anything, like a property's HOA fees, and it answers from the deal's context. 4. Autopilot: Switch Autopilot on and the AI agent runs the conversation itself: it maps out possible conversation paths, reasons about which one moves the deal forward, sends curated listings with photos, then waits for the client and adapts to what they say. 5. Copilot & deal insights: In chat mode the human stays in control and Oracia works as a copilot: suggested replies, a live deal score with sentiment and key moments, and actions like sending a listing's photos, which it only does after the agent confirms. 6. Re-engaging cold leads: Leads that went quiet aren't lost. A cold-lead view shows who needs a nudge, Oracia drafts a personalised follow-up for each, and agents schedule re-engagement plans with attempts, timing and templates. ### Full Stack Developer, GoVibe.live (Macaron Ventures), Oct 2024 – Mar 2025 Remote · New Delhi, India - Led development of GoVibe, a social platform, ensuring a stable launch for 1,000+ users. - Built a microservices backend with Node.js, Kafka, Redis and MySQL for real-time chat. - Engineered a recommendation engine that increased session duration by 30%. - Stack: Node.js, Kafka, Redis, MySQL, React ### Technical Executive, ACM TIET (Association for Computing Machinery), Sep 2023 – Jan 2026 Patiala, India - Led communications for Hacklipse 4.0, a national hackathon with 500+ participants and a ₹2L prize pool. - Managed logistics, sponsorships and a 15-member volunteer team; built the event site with Django. - Formulated 5 AI/ML problem statements with 95% positive feedback. - Stack: Python, Django, AI/ML, Leadership ## Competitive programming (live) - Codeforces (@Naman_18byte): rating 1462, Specialist. Max 1462 · 16 contests. https://codeforces.com/profile/Naman_18byte - LeetCode (@naman_18byte): rating 1832, Top 7.1%. 537 solved · 11 contests. https://leetcode.com/u/naman_18byte - CodeChef (@naman_18byte): rating 1524, 2★. Highest 1524 · 14 contests. https://www.codechef.com/users/naman_18byte - LeetCode: 537 problems solved, 621 submissions and 130 active days in the last year ## Achievements - Smart India Hackathon 2024 Finalist (Top 2.7% nationally): Selected among 1,300 finalist teams out of 49,000+ participants for an AI-based solution. - Merit-Based Scholarship (CGPA 9.14 / 10): Recognized for top academic performance in the branch merit list at Thapar Institute. - Hacklipse 4.0 Lead (500+ participants): Led a national hackathon with a ₹2 Lakh prize pool, a 15-member volunteer team and sponsor outreach.