Skip to content
View romulofff's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report romulofff

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
romulofff/README.md

Rômulo Férrer Filho

Data/ML Engineer at Lexter, building AI for Brazilian social-security law. Fortaleza, CE 🇧🇷

I work at the seam between the model and the product: a prediction has to be well calibrated and it has to survive contact with production.

What I'm working on

  • Case outcome prediction — models that estimate the probability of success for social-security benefit claims (BPC/LOAS), from training to serving: XGBoost, Optuna, SHAP, probability calibration, drift monitoring, and CI gates that stop a bad model before it reaches production.
  • Legal document pipelines — turning court records (PJE) and medical reports into structured features at scale: crawling, LLM extraction, evaluation rubrics, and the data lake underneath.
  • LLM agents that get graded — a medical-examination interview simulator, plus the automated evaluators and eval harnesses that keep it honest.
  • On the side — freelance work on real-time voice agents: speech pipelines, configuration tracking, and metrics for conversations you can't unit-test.

Stack

  • Languages — Python · TypeScript · SQL
  • ML — scikit-learn · XGBoost · SHAP · Optuna · pandas · probability calibration, drift monitoring
  • LLM — LiteLLM · Claude / Anthropic API · MCP · RAG over pgvector · structured extraction · eval harnesses and LLM-as-judge · Langfuse
  • Voice agents — LiveKit Agents · Deepgram · ElevenLabs / Cartesia · Silero VAD · turn detection
  • Backend & infra — FastAPI · PostgreSQL · GCP (Cloud Run, BigQuery, GCS) · Terraform · Docker

Personal projects

  • Nihongo Cards — a Japanese flashcard PWA with FSRS-6 spaced repetition, offline-first, built because I wanted a better way to study kanji.
  • gym-hero — a reinforcement learning environment based on Guitar Hero.
  • Boteco Survivors — a bullet-heaven game set in a Brazilian bar, playable in the browser (source). TypeScript and Canvas, no runtime dependencies. Games are still the reason I got into any of this.

Before this

M.Sc. in Computer Science (Federal University of Ceará, 2024) — multimodal deep reinforcement learning: agents that learn to act from both what they see and what they hear.

B.Sc. in Computer Engineering (UFC, 2021).

Elsewhere

LinkedIn Gmail Website Twitter

Most of what I build these days lives in private repos, so the graph below undersells it.

Rômulo's GitHub stats

Pinned Loading

  1. Agents_that_Listen Agents_that_Listen Public

    Forked from hegde95/Agents_that_Listen

    Train an agent to play VizDoom with multi sensory inputs. Trained using sample factory

    Python

  2. gym-hero gym-hero Public

    Deep Reinforcement Learning Environment Based on Guitar Hero

    Python 2

  3. BoardGameGeekAnalysis BoardGameGeekAnalysis Public

    Exploratory analysis of BoardGameGeek's fórum dataset.

    HTML 2

  4. JoguinhosGratis JoguinhosGratis Public

    This Telegram Bot was developed to send Game Deals to a Telegram Channel

    Python 1 1

  5. datascience_trabalho_sefaz datascience_trabalho_sefaz Public

    Jupyter Notebook 3 1

  6. Visudados_projeto Visudados_projeto Public

    Forked from LauraMMelo/Visudados_projeto

    Contains code to generate interactive visualizations about Ceará Child Death Rate

    JavaScript