Shreeja Singh
Available for Work

Building Intelligent AI Systems for the future

I build production AI agent systems and the evaluation harnesses that catch them when they quietly stop working.

Claude Certified Developer – Foundations
Issued by Anthropic · Verify credential ↗
Shreeja Singh
👩‍🚀

AI Engineer & Multi-Agent Systems Specialist

Agentic AILLM OrchestrationRAG PipelinesModel Evaluation

Hello! I'm Shreeja, an MS Computer Science candidate at Illinois Institute of Technology (GPA 3.6/4.0, graduating May 2027) with 1+ year of industry experience building AI systems that run in production. Currently an AI Engineer Co-op at Aion Labs, where I've shipped 16+ peer-reviewed production PRs across a 3-repo AI data platform — fixing production defects in a Three.js/React 3D engine, building data-quality tools that now gate pipeline releases, and running the platform's first live data refreshes in 75+ days. I'm also a Claude Certified Developer (Anthropic).

Previously a Software Engineer (Quality) at Microchip Technology, where I built an NLP-powered chatbot integrated across ASPICE, JIRA, POLARION and CAD that cut manual search time by 40% for a 50-engineer team.

What I care most about is whether systems actually work. My data-quality checker caught a million-fold unit error in production data on its very first run; on a multi-agent project, I found my own LLM judge was scoring leniently and masking real regressions. Catching what everyone else assumed was fine is the work I find most satisfying.

3+
Projects shipped
1+
Years of experience
16+
Production PRs shipped
5
Dataset families refreshed

Education and Work Experience

Aion Labs
Chicago, IL
Work Experience May 2026 – Present
Microchip Technology
Bengaluru, India
Work Experience Feb 2024 – Jul 2025
Aion Labs
AI Engineer Co-op
📍 Chicago, IL 📅 May 2026 – Present
Key Achievements
  • Shipped 16+ peer-reviewed production PRs across a 3-repo AI data platform: a Python data pipeline (370 tracked artifacts) and a Next.js/TypeScript/Three.js 3D front end, spanning engine fixes, data tooling, and live data refreshes.
  • Diagnosed and fixed 5 production defects in a Three.js/React 3D engine (camera controls, LOD visibility thresholds, raycast hit priority, an infinite render-loop crash) and built its camera fly-to feature, restoring a parked engine to promotion-ready state.
  • Built 3 unit-tested Python data-quality tools (staleness reporting, config-driven quality checks, release reconciliation) now gating pipeline releases. The checker caught a million-fold unit error in production data on its first run, which I traced upstream and fixed with a guarded parser plus regression tests.
  • Ran the platform's first live data refreshes in 75+ days across 5 dataset families (SEC filings, LBNL grid queues, ERCOT PDF reports, token-price APIs), cutting stale artifacts 371 → 252; added a pure-Python PDF-extraction fallback fixing a recurring runner failure.

Skills and Technologies

🤖
Agentic AI & LLMs
Model Context Protocol (MCP)
Claude API
AWS Bedrock
OpenClaw
LLaMA 4
Multi-Agent Orchestration
RAG & Vector Retrieval
Model Routing & Cost Optimization
🔍
Evaluation & Reliability
LLM-as-Judge
RLHF-style Feedback Loops
Benchmark Design
Bias Diagnosis
Cost & Latency Telemetry
⚙️
Backend & APIs
Python
Java
FastAPI
Flask
RESTful API Design
JavaScript
C
🎨
Frontend
Next.js
React
TypeScript
Tailwind
shadcn/ui
Recharts
Three.js
HTML / CSS · Responsive UI
🧠
ML & Data
TensorFlow
PyTorch
PySpark
Spark MLlib
Scikit-learn
Hugging Face
Pandas
NumPy
Feature Engineering · Gradient Boosting
☁️
Cloud & DevOps
AWS: Bedrock · SageMaker · EMR
AWS: Athena · EC2 · IAM · VPC
Docker
Linux
GitHub Actions CI/CD
PII Protection · Least-Privilege IAM
🗄️
Databases
ChromaDB
PostgreSQL
MySQL
SQL
Vector Databases
📋
Process & Standards
JIRA
Confluence
Agile / Scrum / SDLC
ASPICE
ISO 26262
ISO 21434
Git

My Portfolio Highlights

🤖
Jan 2026 – Mar 2026

Autonomous Multi-Agent AI System

An 8-agent routing graph with LLM intent classification directing queries across 4 specialist paths (research, teaching, coding, comparison), grounded by a ChromaDB vector store and real-time web search, deployed through GitHub Actions CI/CD. I built an LLM-as-judge evaluation framework over 5 benchmarked query types reaching 8.0/10 quality at 4.72s latency — then found the judge was scoring leniently and hiding real regressions. Diagnosing and fixing that bias was the most valuable thing I did on the project.

Python FastAPI LLaMA 4 Scout ChromaDB Docker GitHub Actions
↗
📈
Aug 2026 – In progress

Stock Research Agent on MCP

A stock research assistant built on the Model Context Protocol. Rather than one monolithic script, it's a three-server architecture — market data, research, analysis — with a single host holding a dedicated client per server, following the protocol's one-to-one client/server model. The market-data server ships with typed tool schemas auto-derived from Pydantic field definitions via FastMCP, validated in MCP Inspector. Retrieval and evaluation layers are in active development.

Python MCP FastMCP Pydantic yfinance
↗
✈️
Apr 2026 – May 2026

Flight Delay Prediction at Scale

With a five-person team, a 7-phase distributed PySpark pipeline over 6.96M records: cleaning, stratified splits, 8 engineered features, Parquet persistence, profiled with Athena on AWS EMR. I owned the TensorFlow DNN track — and gradient boosting won, at 84.67% accuracy and 79.14% F1 against the DNN's 81.45%, training 3.3× faster. Worth knowing when the simpler model is the right call.

PySpark Spark MLlib TensorFlow AWS SageMaker AWS EMR Athena

Contact Me

Reach out through any of the platforms below. I'm graduating in May 2027 and open to AI/ML and software engineering roles.