AI/ML Engineering Manager · Architect and Builder .
Building and scaling trustworthy, production-grade AI/ML Products. — AI/ML, Agentic AI, LLMs, LLM safety & guardrails, evaluation systems, and scalable ML infrastructure (MLOps / AgentOps).
🔭 Currently: how to build, evaluate, calibrate, Scale and deploy Safely LLM applications and AI agents efficiently. 💞️ Open to collaborating on open-source in AI/ML, LLMs · Generative AI · agentic workflows · AI evaluation · AgenticOps.
📫 GitHub · Pronouns: He/Him
An MLflow-integrated control plane for trusted agentic workflows: governed prompts, tools, skills, MCP servers, workflows, knowledge bases, evaluations, calibration, observability, deployment, and the Aria copilot. The current release tracks MLflow 3.14 and PostgreSQL 17, with a layered architecture, 16 UI cookbooks, and a published documentation site.
AI platform · MLflow · AgentOps · governance · evaluation · observability
Safety benchmark gains do not guarantee safety transfer. A paired, same-checkpoint study of how compact prompt-safety guards specialize, transfer, and compose — LoRA-SFT vs. base-anchored KL-SFT across four instruction checkpoints, a dual-labeled mortgage benchmark, and an analysis-preregistered panel of released vendor guards. The v0.0.1 research hub brings together the published HTML report and the unified report PDF. The release records explicit study state, evidence tiers, verification paths, and redistribution decisions; 28 of 32 generated inputs verify in the standard environment, with the remaining four requiring the pinned analysis environment.
LLM safety · guard models · LoRA / SFT · preregistration · benchmark transfer · reproducible research · research release
Traces establish recurrence, not admissibility. A trace-to-program compiler that turns repeated read-only agent prefixes into deterministic guarded programs — and refuses whenever the evidence cannot license one. Typed value provenance, effect and position barriers, a bounded 23-operator DSL, runtime verification, and a finite-sample selective-risk gate whose default output is retirement. Across three live-provider GitHub workflow families it matches the baseline on 90/90 exact outcomes (versus 89/90), while reducing provider requests by 66.6%; on NESTFUL and API-Bank every recurrent family retires — which is the result, not a failure. The newly published artifact shelf documents the evidence and its limits.
agent optimization · program synthesis · provenance · selective risk control · LLM agents · reproducible research
Faithful ICLR 2026 implementation — evolving, self-improving context playbooks for LLM agents via a Generator → Reflector → Curator loop with incremental delta updates. OpenAI Agents SDK support, 163 tests, and an 11-recipe cookbook.
context engineering · self-improving agents · in-context learning · agent memory
A clean reference implementation of the Agent-to-Agent (A2A) protocol — specialized AI agents that discover each other and collaborate over JSON-RPC 2.0. Python · FastAPI · Pydantic, with 147 tests at 93% coverage and a worked cookbook.
multi-agent systems · agent interoperability · A2A · FastAPI
Enterprise compliance automation — turn compliance documents into queryable knowledge graphs via a multi-agent AI pipeline, with an interactive graph explorer.
knowledge graphs · compliance · RegTech · multi-agent · JanusGraph
Reinforcement learning (PPO/A2C/DQN) that dynamically tunes AML risk-scoring weights per case — Gymnasium env, FastAPI backend, React training dashboard.
reinforcement learning · AML · RegTech · PPO · risk scoring
Domain-agnostic, single-node, docker compose-portable tabular AutoML pipeline (Dagster + FLAML + MLflow + FastAPI) with drift monitoring, an online feature store, a Streamlit dashboard, and one-command start/stop. The real-estate buyer-stage classifier ships as the worked example — all domain knowledge lives in YAML, never in framework code.
AutoML · MLflow · Dagster · drift detection · tabular ML
AI-powered mock-interview assistant for ML engineering, leadership/behavioural, and coding interviews. Gradio 6 · LangChain 1.x · OpenAI · Whisper.
interview prep · LangChain · speech-to-text · generative AI
Agentic AI · LLM safety & guardrails · prompt-injection / jailbreak detection · LLM & agent evaluation · program synthesis & selective risk control · MLflow / LLMOps / MLOps · RAG · reinforcement learning (RLHF/GRPO/PPO) · fine-tuning (SFT/LoRA) · knowledge graphs · multi-agent systems
Python · PyTorch · Transformers · TRL · MLflow · FastAPI · LangChain · React
⭐ If any of these are useful, a star helps others find them too.



