AI ENGINEER

Building intelligent systems
that reason, retrieve,
and act.

I design and deploy production AI systems using LLMs, agentic workflows, retrieval-augmented generation, semantic search, and scalable AI infrastructure.

Agent architecture
User
Agent Orchestrator
Planner / Router
Specialized Agents
Tools + Retrieval
LLM · GPT-4
Response
01About

I build AI systems beyond simple LLM wrappers.

I build AI systems beyond standalone model calls. My work spans orchestration, retrieval, memory, tool use, knowledge systems, APIs, evaluation, and deployment.

I focus on turning LLM capabilities into production workflows by treating retrieval quality, system state, deployment, and reliability as first-class engineering problems.

Agentic Systems

Multi-agent orchestration with LangGraph, stateful workflows, and tool-using agents.

RAG & Retrieval

Vector and graph retrieval with FAISS, ChromaDB, Pinecone, and Neo4j for grounded responses.

LLM Engineering

Prompt engineering and model workflows across GPT-4, LLaMA2, and T5 for domain-specific automation.

Production AI / MLOps

Shipping AI services with Python APIs, Docker, Kubernetes, AWS, and CI/CD workflows.

0%

Faster service response

Service response

0%

Higher ingestion throughput

Data ingestion

0%

Lower inference latency

Inference

0%

Reduced manual review / triage

Claims review

02Projects

Featured AI systems

Two production-focused case studies. The diagrams are representative architectures based on the technologies and system components used, rather than literal infrastructure maps. Hover a node for context.

Case Study 01

Enterprise Multi-Agent RAG System

Problem

Healthcare claims review required substantial manual effort, with fraud evidence distributed across documents, metadata, and enterprise knowledge sources.

System

A representative architecture based on the technologies and system components used: LangGraph coordinates retrieval and reasoning workflows, vector stores provide grounded context, and GPT-4 / LLaMA2 process retrieved evidence for claims decision support.

LangGraphLangChainGPT-4LLaMA2FAISSChromaDBPythonRAG
35%

Reduction in manual review time

22%

Higher fraud detection accuracy

Representative architecture

Animated to illustrate a request progressing through the major system stages.

Claim / Query
LangGraph Orchestrator
Retrieval
Reasoning
Evidence
FAISS / ChromaDB
GPT-4 / LLaMA2
Metadata
Context Aggregation
Decision Support
Case Study 02

Production Agentic AI Assistant

Problem

Enterprise workflows needed an assistant that could preserve context, use tools, retrieve grounded knowledge, and support multi-step interactions instead of relying on isolated model calls.

System

A representative architecture based on my agentic AI work: LangGraph manages workflow state, memory preserves context, ReAct-style tool use handles actions, and vector / graph retrieval supplies relevant enterprise knowledge before the final reasoning step.

LangGraphLangChain ReActGPT-4Neo4jVector DatabasesPythonTool-using Agents
30%

Fewer SLA violations

Representative architecture

Animated to illustrate a request progressing through the major system stages.

User Request
LangGraph Orchestrator
ReAct Agent
Memory
Tools
Retrieval
Workflow State
Enterprise Actions
Vector DB / Neo4j
LLM Reasoning
Final Response
03Experience

Career timeline

A curated view of the work most relevant to Generative AI, RAG, agentic systems, and production ML.

Responsive

Jul 2024 — Present

AI Engineer · Remote

  • Designed LangGraph-powered GPT-4 chatbots for internal automation, improving service response time by 35% through multi-agent orchestration.
  • Built memory-aware, tool-using agents with the LangChain React Agent framework for multi-step reasoning workflows.
  • Engineered RAG pipelines with FAISS, ChromaDB, and Pinecone, increasing ingestion throughput 40% through token-optimized chunking and extending retrieval with Neo4j graph-based Q&A workflows.
LangGraphLangChainGPT-4FAISSChromaDBNeo4jPinecone

Wipro

Dec 2021 — Mar 2024

GenAI / ML Engineer · Hyderabad, India

  • Designed RAG systems integrating GPT-4, LLaMA2, and T5 for healthcare claims automation and document summarization.
  • Reduced manual triage by 35% through optimized RAG workflows and metadata retrieval for fraud detection.
  • Reduced inference latency 20% through deployment optimization on AWS Lambda, EC2, and S3.
LangChainFAISSChromaDBGPT-4LLaMA2T5AWS SageMaker

Zensar Technologies

Dec 2020 — Nov 2021

Associate AI Engineer · Hyderabad, India

  • Deployed transformer-based NER and classification models on AWS and Azure for finance and telecom.
  • Built automated ETL pipelines with Airflow and Spark for large-scale structured and unstructured data.
PythonTensorFlowPyTorchHugging FaceAirflowAWSAzure
04Stack

Technical stack

Agentic AI

LangGraphLangChainReAct AgentsTool-using AgentsMulti-Agent Systems

LLM & RAG

GPT-4LLaMA2RAGPrompt EngineeringHugging FaceSemantic Search

Retrieval & Knowledge

FAISSChromaDBPineconeNeo4jSQL

AI / ML

PythonPyTorchTensorFlowScikit-learnBERTSBERT

Production AI

AWSSageMakerDockerKubernetesOpenShiftCI/CDREST APIs
05Education & Certifications

Education

M.S. in Computer Science

Sacred Heart University

Fairfield, Connecticut

Jan 2024 — Mar 2025·CGPA 3.76 / 4

B.Tech in Electronics and Communication Engineering

Lakireddy Bali Reddy College of Engineering

India

Dec 2021·CGPA 7.59

Certifications

Generative AI LLMs

NVIDIA-Certified Associate

Generative AI with LLMs

DeepLearning.AI (Coursera)

2025

Artificial Intelligence A-Z

Udemy

2025

Certified Machine Learning Engineer

Udemy

2024

06 / CONTACT

Let's build intelligent systems.

I'm interested in opportunities involving Generative AI, Agentic AI, RAG, LLM infrastructure, and production AI systems.

nune@ai-engineer:~
>whoami
Sai Kumar Nune
AI Engineer
>focus
>architecture