Skip to content
View YuliyaCarvalho's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report YuliyaCarvalho

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
YuliyaCarvalho/README.md
 ╔══════════════════════════════════════════════════════════════╗
 ║   YULIA CARVALHO  ·  DATA ANALYST & DATA ENGINEER            ║
 ║   Lisbon, Portugal  ·  Available · Open to Remote / On-site  ║
 ╚══════════════════════════════════════════════════════════════╝

Portfolio LinkedIn Kaggle ORCID


A scientist's analyst.

I spent fifteen years in high-impact research environments - molecular biology, biomedical science, large-scale diagnostics - where the underlying work was always the same: design a rigorous process, gather imperfect data, extract reliable signal from noise. That instinct transferred cleanly into data engineering and analytics.

What I enjoy most is what I call forensic EDA: a structured, investigative approach to exploratory analysis. Analysis, for me, starts before dashboards or models — with questioning the integrity of the data itself. Cleaning inconsistencies, validating assumptions, identifying anomalies, tracing duplicates, resolving cross-field conflicts. I approach datasets the way I once approached laboratory experiments: carefully, skeptically, and methodically, until the underlying logic becomes clear.

I'm drawn to the engineering side of analytics: building reproducible pipelines, defining clear data contracts, designing layered transformation architectures, and documenting logic so that analytical outputs can actually be trusted downstream.

15+ years research · 12 peer-reviewed publications · Lisbon, PT
EN · PT · RU


Selected Projects

🌦 IPMA Weather Pipeline — Automated Ingestion, Transformation & Monitoring

Python BigQuery dbt Core GitHub Actions Looker Studio

End-to-end data engineering pipeline ingesting hourly weather observations from 222 meteorological stations across Portugal via the IPMA public API. Implements a layered architecture (raw → staging → mart) with automated data quality testing, incremental batch loading, and daily scheduling.

  • Append-only raw layer with ingested_at timestamps for full auditability and root-cause tracing
  • Dual-table ingestion pattern reflecting different update semantics (append vs. truncate-reload)
  • Two-layer dbt transformation: type casting, sentinel value nullification, wind direction decoding, deduplication on natural key
  • 8 automated dbt tests covering null constraints, timestamp validity, and measurement field ranges — pipeline fails automatically on violation
  • Full pipeline orchestrated via GitHub Actions cron with failure notifications and run logs

~4 151 observations/day · 222 stations · 8 quality tests

→ Repository


📦 Olist: Retention, Logistics & Risk — E-Commerce Operational Analytics

SQL BigQuery R Kaggle

Capstone project for the Google Data Analytics Certificate. A comprehensive operational study of Brazil's Olist marketplace diagnosing the structural drivers of customer dissatisfaction, logistics failure, and seller performance risk.

  • End-to-end pipeline connecting BigQuery to a statistical environment, integrating transaction, logistics, and sentiment data
  • Multiple linear regression isolating delivery delay as the dominant predictor of customer satisfaction
  • Composite seller risk scoring system weighing historical lateness, delay severity, and business value at risk
  • Origin–destination route analysis distinguishing individual seller inefficiencies from systemic regional constraints
  • Interaction regression measuring how seasonal shopping events amplify the impact of delivery delays

100K+ orders · 18 business questions · 1 operational dashboard

→ Notebook · → Repository · → Dashboard


🌿 Biodiversity Loss in Europe — Environmental Analytics

Python SQL Google Colab Looker Studio

Graduation project for the Le Wagon Data Analytics Bootcamp. Structured environmental analysis examining the relationship between land-use patterns, protected-area coverage, and the Biodiversity Intactness Index across 27 EU countries.

  • Multi-source integration: land-use, protected-area, and biodiversity datasets into a unified analytical framework
  • Cross-country structural pattern comparison identifying recurring drivers of biodiversity outcomes
  • Looker Studio dashboards translating findings into accessible, decision-oriented visual communication

27 EU countries · 5 data sources · team graduation project

→ Dashboard / Report


Toolkit

Languages SQL — CTEs, window functions, validation logic · Python — data analysis, pipeline scripting · R — statistical modeling, reporting · HTML5 / CSS3

Data Engineering BigQuery · dbt Core (raw → staging → mart) · GitHub Actions · Airflow (concepts, DAGs) · Fivetran (ELT concepts) · DuckDB

Analysis Libraries pandas · NumPy · scikit-learn · Prophet · tidyverse / lubridate · BeautifulSoup

Visualization & BI Looker Studio · Power BI (DAX, data modeling) · matplotlib / seaborn · Plotly · ggplot2

Workflow Git / GitHub · Jupyter / Colab · VS Code / RStudio · Kaggle · Bash / CLI


Background

Period Role Organization
2025 – present Independent Data Analyst Self-directed — pipeline development, BI, applied upskilling
2023 – 2025 Laboratory Data Manager GIMM – Gulbenkian Institute for Molecular Medicine
2020 – 2023 COVID-19 Diagnostic Specialist iMM – Instituto de Medicina Molecular (200K+ tests coordinated)
2018 – 2020 Research Fellow iMM – Instituto de Medicina Molecular
2012 – 2017 Scientific Researcher / PhD Student Goethe University, Frankfurt
2011 – 2012 Invited Researcher Max Planck Institute for Heart and Lung Research

Certificates

  • Google Data Analytics Professional Certificate · Google / Coursera · Jan 2026
  • Data Analytics Bootcamp (DGERT-certified) · Le Wagon · Nov 2025
  • PRINCE2 Foundation · PeopleCert / AXELOS · Jan 2024
  • ETL in Python and SQL · LinkedIn Learning · Feb 2026

Publications — Selected

12 peer-reviewed papers across cardiovascular biology, molecular biology, and bioinformatics.

  • Cardiovascular Research (2024) · Developmental Cell (2022) · PNAS (2021)
  • Circulation (2017) · Circulation Research (2014, 2018) · Nucleic Acids Research (2012)

→ Full list on ORCID · → Full list on Portfolio


"Turning chaos into structured clarity is not just a task — it is the part of the process I enjoy most."

yuliya.carvalho@gmail.com · Portfolio · LinkedIn

Popular repositories Loading

  1. lab-github-practice lab-github-practice Public

    Forked from ironhack-labs/lab-github-practice

    This is a repository for learning about GitHub features such as cloning, branching, and pull requests.

    HTML

  2. lab-git-practice-data-ai lab-git-practice-data-ai Public

    Forked from ironhack-labs/lab-git-practice-data-ai

    Jupyter Notebook

  3. LeWagon_ToyStore LeWagon_ToyStore Public

    LeWagon DA Bootcamp: week 3_day 4_challenge2: GTM_setup

    HTML

  4. dbt-delivers dbt-delivers Public

    This is a practice repository for transforming a dataset on Food Delivery in France in dbt cloud

    HTML

  5. YuliyaCarvalho YuliyaCarvalho Public

  6. Portugal-Property-Atlas Portugal-Property-Atlas Public

    Jupyter Notebook