╔══════════════════════════════════════════════════════════════╗
║ YULIA CARVALHO · DATA ANALYST & DATA ENGINEER ║
║ Lisbon, Portugal · Available · Open to Remote / On-site ║
╚══════════════════════════════════════════════════════════════╝
I spent fifteen years in high-impact research environments - molecular biology, biomedical science, large-scale diagnostics - where the underlying work was always the same: design a rigorous process, gather imperfect data, extract reliable signal from noise. That instinct transferred cleanly into data engineering and analytics.
What I enjoy most is what I call forensic EDA: a structured, investigative approach to exploratory analysis. Analysis, for me, starts before dashboards or models — with questioning the integrity of the data itself. Cleaning inconsistencies, validating assumptions, identifying anomalies, tracing duplicates, resolving cross-field conflicts. I approach datasets the way I once approached laboratory experiments: carefully, skeptically, and methodically, until the underlying logic becomes clear.
I'm drawn to the engineering side of analytics: building reproducible pipelines, defining clear data contracts, designing layered transformation architectures, and documenting logic so that analytical outputs can actually be trusted downstream.
15+ years research · 12 peer-reviewed publications · Lisbon, PT
EN · PT · RU
Python BigQuery dbt Core GitHub Actions Looker Studio
End-to-end data engineering pipeline ingesting hourly weather observations from 222 meteorological stations across Portugal via the IPMA public API. Implements a layered architecture (raw → staging → mart) with automated data quality testing, incremental batch loading, and daily scheduling.
- Append-only raw layer with
ingested_attimestamps for full auditability and root-cause tracing - Dual-table ingestion pattern reflecting different update semantics (append vs. truncate-reload)
- Two-layer dbt transformation: type casting, sentinel value nullification, wind direction decoding, deduplication on natural key
- 8 automated dbt tests covering null constraints, timestamp validity, and measurement field ranges — pipeline fails automatically on violation
- Full pipeline orchestrated via GitHub Actions cron with failure notifications and run logs
~4 151 observations/day · 222 stations · 8 quality tests
SQL BigQuery R Kaggle
Capstone project for the Google Data Analytics Certificate. A comprehensive operational study of Brazil's Olist marketplace diagnosing the structural drivers of customer dissatisfaction, logistics failure, and seller performance risk.
- End-to-end pipeline connecting BigQuery to a statistical environment, integrating transaction, logistics, and sentiment data
- Multiple linear regression isolating delivery delay as the dominant predictor of customer satisfaction
- Composite seller risk scoring system weighing historical lateness, delay severity, and business value at risk
- Origin–destination route analysis distinguishing individual seller inefficiencies from systemic regional constraints
- Interaction regression measuring how seasonal shopping events amplify the impact of delivery delays
100K+ orders · 18 business questions · 1 operational dashboard
→ Notebook · → Repository · → Dashboard
Python SQL Google Colab Looker Studio
Graduation project for the Le Wagon Data Analytics Bootcamp. Structured environmental analysis examining the relationship between land-use patterns, protected-area coverage, and the Biodiversity Intactness Index across 27 EU countries.
- Multi-source integration: land-use, protected-area, and biodiversity datasets into a unified analytical framework
- Cross-country structural pattern comparison identifying recurring drivers of biodiversity outcomes
- Looker Studio dashboards translating findings into accessible, decision-oriented visual communication
27 EU countries · 5 data sources · team graduation project
Languages
SQL — CTEs, window functions, validation logic · Python — data analysis, pipeline scripting · R — statistical modeling, reporting · HTML5 / CSS3
Data Engineering
BigQuery · dbt Core (raw → staging → mart) · GitHub Actions · Airflow (concepts, DAGs) · Fivetran (ELT concepts) · DuckDB
Analysis Libraries
pandas · NumPy · scikit-learn · Prophet · tidyverse / lubridate · BeautifulSoup
Visualization & BI
Looker Studio · Power BI (DAX, data modeling) · matplotlib / seaborn · Plotly · ggplot2
Workflow
Git / GitHub · Jupyter / Colab · VS Code / RStudio · Kaggle · Bash / CLI
| Period | Role | Organization |
|---|---|---|
| 2025 – present | Independent Data Analyst | Self-directed — pipeline development, BI, applied upskilling |
| 2023 – 2025 | Laboratory Data Manager | GIMM – Gulbenkian Institute for Molecular Medicine |
| 2020 – 2023 | COVID-19 Diagnostic Specialist | iMM – Instituto de Medicina Molecular (200K+ tests coordinated) |
| 2018 – 2020 | Research Fellow | iMM – Instituto de Medicina Molecular |
| 2012 – 2017 | Scientific Researcher / PhD Student | Goethe University, Frankfurt |
| 2011 – 2012 | Invited Researcher | Max Planck Institute for Heart and Lung Research |
- Google Data Analytics Professional Certificate · Google / Coursera · Jan 2026
- Data Analytics Bootcamp (DGERT-certified) · Le Wagon · Nov 2025
- PRINCE2 Foundation · PeopleCert / AXELOS · Jan 2024
- ETL in Python and SQL · LinkedIn Learning · Feb 2026
12 peer-reviewed papers across cardiovascular biology, molecular biology, and bioinformatics.
- Cardiovascular Research (2024) · Developmental Cell (2022) · PNAS (2021)
- Circulation (2017) · Circulation Research (2014, 2018) · Nucleic Acids Research (2012)
→ Full list on ORCID · → Full list on Portfolio
"Turning chaos into structured clarity is not just a task — it is the part of the process I enjoy most."