📍 Chicago, IL ·
"Most AI demos break. I build the systems that don't."
I spent 13 plus years building IT infrastructure before I moved into AI. That background shapes how I work: I don't just build things that demo well. I build things that hold up under edge cases, under load, and over time.
Right now I'm focused on LLM systems, RAG pipelines, and the evaluation infrastructure that makes AI actually trustworthy in production.
LLM Systems & Prompt Engineering Few-shot, Chain-of-Thought, and ReAct prompting. RAG pipelines with retrieval, grounding, and structured output. LLM evaluation frameworks, hallucination detection, and benchmarking across GPT-4, Claude, LLaMA, and Gemini.
AI Automation & Systems Design End-to-end workflow automation in Python. Multi-step AI pipelines designed for real-world failure modes. Human-in-the-loop (HITL) architectures. Prompt chaining with structured reasoning.
Data Engineering & Analytics SQL analytics systems in PostgreSQL and SQL Server. ETL pipelines, star schema modeling, KPI design. Cohort analysis, retention modeling, Power BI dashboards with DAX.
Machine Learning Classification, regression, and clustering with scikit-learn. Feature engineering, model evaluation, fraud detection, and churn prediction pipelines.
| What changed | By how much |
|---|---|
| Manual work eliminated | ~70% |
| Hallucination rate reduced | ~40% |
| Pipeline turnaround time | ~60% faster |
| Reporting time saved | ~50% |
| Average F1-Score | 0.85+ |
A production-grade harness for testing LLM prompts systematically ,not by eyeballing output.
- Four-dimensional weighted scoring (JSON structure, required keys, safety compliance, content rules).
- Simulated OpenAI provider for continuous testing without API costs. 100% pass rate on test suite.
- Reduced hallucinations 30–40% through structured prompt optimization.
- Tech:
PythonOpenAI APIPrompt EngineeringLLM Evaluation
Built an AI-powered Prompt Engineering Evaluation System that analyzes, scores, compares, and refines LLM prompts using structured evaluation metrics instead of manual guesswork.
- Designed a multi-dimensional prompt scoring framework measuring clarity, specificity, robustness, efficiency, and alignment, including hallucination-risk analysis and AI-generated optimization suggestions.
- Implemented prompt versioning and head-to-head comparison workflows to track iterative improvements and evaluate prompt performance across different task types (RAG, summarization, reasoning, extraction, etc.). -Developed a full-stack React application integrated with the Anthropic Claude API for real-time prompt evaluation, refinement, and predicted output simulation.
- Tech:
React18JavaScriptPrompt EngineeringLLM workflowsTailwind CSSAnthropic Claude API (Claude Sonnet 4)AI Evaluation Systems
End-to-end analytics pipeline modeling user performance & engagement
- Leaderboard, streak & rolling 7-day metrics
- Question difficulty modeling & query optimization
- Tech:
PostgreSQLAdvanced SQLWindow FunctionsKPI Modeling
E-Commerce Analytics Platform (SQL-Only) Complete SQL-based analytics system for transactional data
- Revenue, AOV, LTV & retention metrics
- Cohort analysis & reusable reporting views
- Tech:
PostgreSQLSQLAggregationsWindow Functions
Customer Churn Prediction End-to-end ML pipeline for churn prediction using telecom data
- Feature engineering, model training & evaluation
- Business-driven retention insights
- Tech:
PythonPandasscikit-learnClassification
Credit Card Fraud Detection Fraud detection model addressing severe class imbalance
- Precision/Recall optimization & F1-score evaluation
- False positive minimization strategy
- Tech:
PythonPandasscikit-learn
NLP Chatbot (Flask App) Rule-based chatbot with TF-IDF & cosine similarity
- REST API backend & preprocessing pipeline
- Tech:
PythonFlaskNLTKHTMLCSS
- Production-grade RAG systems with evaluation scoring
- LLM workflow automation pipelines.
- ML monitoring and drift detection.
- Advanced prompt optimization frameworks.
Full-time remote roles in Generative AI Engineering, Prompt Engineering, AI Systems Engineering, LLM Application Development, and Data Engineering.