Skip to content
View rrahimi-uci's full-sized avatar

Highlights

  • Pro

Block or report rrahimi-uci

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
rrahimi-uci/README.md

Reza Rahimi

AI/ML Engineering Manager · Architect and Builder .

Building and scaling trustworthy, production-grade AI/ML Products. — AI/ML, Agentic AI, LLMs, LLM safety & guardrails, evaluation systems, and scalable ML infrastructure (MLOps / AgentOps).

🔭 Currently: how to build, evaluate, calibrate, Scale and deploy Safely LLM applications and AI agents efficiently. 💞️ Open to collaborating on open-source in AI/ML, LLMs · Generative AI · agentic workflows · AI evaluation · AgenticOps.

📫 GitHub · Pronouns: He/Him


🚀 Featured Projects

An MLflow-integrated control plane for trusted agentic workflows: governed prompts, tools, skills, MCP servers, workflows, knowledge bases, evaluations, calibration, observability, deployment, and the Aria copilot. The current release tracks MLflow 3.14 and PostgreSQL 17, with a layered architecture, 16 UI cookbooks, and a published documentation site. AI platform · MLflow · AgentOps · governance · evaluation · observability

Safety benchmark gains do not guarantee safety transfer. A paired, same-checkpoint study of how compact prompt-safety guards specialize, transfer, and compose — LoRA-SFT vs. base-anchored KL-SFT across four instruction checkpoints, a dual-labeled mortgage benchmark, and an analysis-preregistered panel of released vendor guards. The v0.0.1 research hub brings together the published HTML report and the unified report PDF. The release records explicit study state, evidence tiers, verification paths, and redistribution decisions; 28 of 32 generated inputs verify in the standard environment, with the remaining four requiring the pinned analysis environment. LLM safety · guard models · LoRA / SFT · preregistration · benchmark transfer · reproducible research · research release

Traces establish recurrence, not admissibility. A trace-to-program compiler that turns repeated read-only agent prefixes into deterministic guarded programs — and refuses whenever the evidence cannot license one. Typed value provenance, effect and position barriers, a bounded 23-operator DSL, runtime verification, and a finite-sample selective-risk gate whose default output is retirement. Across three live-provider GitHub workflow families it matches the baseline on 90/90 exact outcomes (versus 89/90), while reducing provider requests by 66.6%; on NESTFUL and API-Bank every recurrent family retires — which is the result, not a failure. The newly published artifact shelf documents the evidence and its limits. agent optimization · program synthesis · provenance · selective risk control · LLM agents · reproducible research

Faithful ICLR 2026 implementation — evolving, self-improving context playbooks for LLM agents via a Generator → Reflector → Curator loop with incremental delta updates. OpenAI Agents SDK support, 163 tests, and an 11-recipe cookbook. context engineering · self-improving agents · in-context learning · agent memory

A clean reference implementation of the Agent-to-Agent (A2A) protocol — specialized AI agents that discover each other and collaborate over JSON-RPC 2.0. Python · FastAPI · Pydantic, with 147 tests at 93% coverage and a worked cookbook. multi-agent systems · agent interoperability · A2A · FastAPI

Enterprise compliance automation — turn compliance documents into queryable knowledge graphs via a multi-agent AI pipeline, with an interactive graph explorer. knowledge graphs · compliance · RegTech · multi-agent · JanusGraph

Reinforcement learning (PPO/A2C/DQN) that dynamically tunes AML risk-scoring weights per case — Gymnasium env, FastAPI backend, React training dashboard. reinforcement learning · AML · RegTech · PPO · risk scoring

Domain-agnostic, single-node, docker compose-portable tabular AutoML pipeline (Dagster + FLAML + MLflow + FastAPI) with drift monitoring, an online feature store, a Streamlit dashboard, and one-command start/stop. The real-estate buyer-stage classifier ships as the worked example — all domain knowledge lives in YAML, never in framework code. AutoML · MLflow · Dagster · drift detection · tabular ML

AI-powered mock-interview assistant for ML engineering, leadership/behavioural, and coding interviews. Gradio 6 · LangChain 1.x · OpenAI · Whisper. interview prep · LangChain · speech-to-text · generative AI


🛠️ Focus Areas

Agentic AI · LLM safety & guardrails · prompt-injection / jailbreak detection · LLM & agent evaluation · program synthesis & selective risk control · MLflow / LLMOps / MLOps · RAG · reinforcement learning (RLHF/GRPO/PPO) · fine-tuning (SFT/LoRA) · knowledge graphs · multi-agent systems

Python · PyTorch · Transformers · TRL · MLflow · FastAPI · LangChain · React


⭐ If any of these are useful, a star helps others find them too.

Pinned Loading

  1. caliber-suite caliber-suite Public

    Open-source MLflow plugin for AI agents and agentic workflows: prompts, tools, skills, MCP servers, RAG knowledge bases, evaluation, deployment, observability, and Aria copilot.

    Python 1

  2. guarded-agentic-compaction guarded-agentic-compaction Public

    Compile the routine, refuse the uncertain — a research library and paper on compiling recurrent read-only regions of tool-using LLM agents into deterministic guarded programs, admitted only when va…

    Python 1

  3. safety-guard-dynamics safety-guard-dynamics Public

    Auditable research on how compact safety guards specialize, transfer, compose, and adapt across objectives, benchmarks, and regulated domains.

    Python 1

  4. agentic-context-engineering agentic-context-engineering Public

    ACE — Agentic Context Engineering: evolving, self-improving context playbooks for LLM agents. Faithful ICLR 2026 implementation with OpenAI Agents SDK support.

    Python 1

  5. policy-to-knowledge policy-to-knowledge Public

    Enterprise compliance automation: transform compliance documents into queryable knowledge graphs via a multi-agent AI pipeline, with an interactive graph explorer.

    HTML 1

  6. rl-anti-money-laundry rl-anti-money-laundry Public

    Reinforcement learning (PPO/A2C/DQN) that dynamically tunes Anti-Money-Laundering risk-scoring weights per case — Gymnasium env, FastAPI backend, and a React training dashboard.

    Python 1