Skip to content
View ceugenia's full-sized avatar

Block or report ceugenia

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
ceugenia/README.md

👋 Hi there, my name is Constanza Eugenia Perez! (AKA ceugenia)

Data Analytics Professional | SnowPro Core Certified | SQL | Python | Snowflake | Bioinformatics

Data professional with a background in microbiology, bioinformatics, and computational research, focused on data analytics and building expertise toward data engineering. I have hands-on experience working with Python, R, SQL, single-cell RNA sequencing data, exploratory data analysis, and computational biology workflows. My research background has given me experience working with large and complex datasets, developing reproducible analytical workflows, troubleshooting technical problems, and communicating data-driven findings. I am SnowPro Core Certified and actively expanding my skills in cloud data platforms, SQL, data warehousing, and modern data engineering practices.

Technical Interests: Data Analytics · SQL · Python · Snowflake · Data Warehousing · ETL/ELT · Exploratory Data Analysis (EDA) · Data Visualization · Bioinformatics · Computational Biology

Career Path: Currently focused on Data Analytics with strategic growth toward Data Engineering roles. Seeking remote Data Analyst, Analytics Engineer, and Bioinformatics/Data Analyst positions where I can leverage scientific background to drive data-driven insights.


🖥️ Current Focus

  • Data analytics & exploration of large-scale single-cell RNA-seq datasets
  • Building interactive data visualizations with Shiny apps for exploratory analysis
  • Developing data quality workflows for transcriptomic data pipelines
  • Statistical analysis & pattern identification in complex biological datasets
  • ETL/ELT fundamentals and data pipeline design (building toward data engineering)

🛠️ Tech Stack

Top Skills: Snowflake · SQL · Python · Data Analytics · Data Visualization · scRNA-seq · Seurat · Scanpy · R

Languages & Environments: Python · R · Bash · SQL · Jupyter · Google Colab

Data Analytics & Visualization: Exploratory Data Analysis (EDA) · Data Wrangling · Statistical Analysis · Plotly · Dash · Shiny · ggplot2 · Data Storytelling

Databases & Platforms: SQL · Snowflake · Data Curation · Data Cleaning · Data Quality · Database Management · Electronic Data Capture (EDC)

Programming: Python (Pandas, NumPy, Scikit-learn) · R (Tidyverse, Seurat, Shiny) · MATLAB · Java · C/C++

Data Analysis & Machine Learning: Decision Trees · Random Forests · XGBoost · Logistic Regression · Model Evaluation · Feature Engineering

Bioinformatics & Computational Biology: Single-Cell RNA-seq · Bulk RNA-seq · Transcriptomic Profiling · Seurat · Scanpy · CIARA · Monocle3 · Destiny

Visualization & Development: Plotly · Dash · Shiny · ImageJ/Fiji · OpenCV · scikit-image · Git · GitHub · Linux

Professional: Data Quality Control · Documentation · SOP Development · Research Collaboration · HIPAA Compliance

Languages: English (Native or Bilingual) · Spanish (Native or Bilingual)


📁 Featured Projects

🧬 immune-chord

Data analytics pipeline for identifying rare cell populations in single-cell RNA-seq data

  • Developed an end-to-end analytical workflow for processing and exploring single-cell RNA sequencing datasets
  • Implemented data quality control, normalization, clustering, PCA, UMAP, and population-level analysis using Seurat
  • Created exploratory data analysis (EDA) visualizations to identify biologically meaningful patterns
  • Organized reproducible analytical workflows using Conda environment management
  • Applied data validation practices to ensure data integrity and reproducibility

Outcome: End-to-end reproducible analytics workflow with publication-ready visualizations
Tech: R · Seurat · Python · Data Analytics · UMAP · EDA · Quality Control


🖼️ PTBP1 ImageJ Analysis Suite

Data analytics pipeline for quantifying protein localization from microscopy datasets

  • Designed and automated an analytical pipeline to explore subcellular PTBP1 localization patterns
  • Processed and analyzed 26,000+ cell images using Python data analysis scripts
  • Applied statistical analysis and feature extraction using OpenCV and scikit-image
  • Developed data visualizations and analytics reports to communicate findings

Outcome: Analyzed 26,040 cells; identified novel localization patterns; presented at CSUF Research Symposium (2021)
Tech: Python · ImageJ/Fiji · OpenCV · scikit-image · Dash · Data Analytics · Statistical Analysis


🔬 CIARA-MGCrestLab-SC 🔒 (Private)

Comparative data analytics: Benchmarking rare-cell detection algorithms on human gastrula datasets

  • Performed exploratory data analysis (EDA) on single-cell RNA-seq datasets using R and Python
  • Applied data preprocessing, normalization, dimensionality reduction, and clustering workflows
  • Developed interactive R Shiny application for exploring and visualizing gene-expression patterns
  • Conducted statistical analysis for pseudotime and developmental trajectory data interpretation
  • Created data visualizations to compare algorithm performance and identify rare cell populations

Outcome: CIARA achieved recall = 0.92 at low read depth; identified rare hemogenic endothelial progenitors
Tech: R · Python · Seurat · SQL · Data Analytics · Visualization · Statistical Analysis


🧪 MGCrestLab ShinyApp 🔒 (Private)

Interactive data analytics dashboard for single-cell RNA-seq exploratory analysis

  • Designed interactive data analytics interface for exploring Seurat clusters and gene expression patterns
  • Developed R Shiny applications for visual data exploration and discovery
  • Performed data transformation, quality checks, and statistical analysis on transcriptomic datasets
  • Created data visualizations (heatmaps, scatter plots, bar charts) for biological interpretation
  • Reduced data exploration time by ~40% for research collaborators through interactive analytics

Outcome: Significantly improved data exploration efficiency and accessibility
Tech: R · Shiny · Seurat · ggplot2 · Plotly · VennDiagram · Data Analytics


👩‍💻 Experience

Computational Biology Lead | MGCrest Lab, University of California Riverside

November 2023 – Present (2 years 11 months)

Single cell transcriptomic analysis and data analytics of neural crest origin

  • Lead data analysis and exploration workflows within the lab, providing analytical support and insights to lab members
  • Design and maintain analytical workflows for large-scale single-cell and bulk transcriptomic datasets, including data ingestion, cleaning, quality control, transformation, and exploration
  • Develop reproducible analytical pipelines using Python, R, SQL, Seurat, Scanpy, and Linux
  • Perform exploratory data analysis (EDA), data validation, preprocessing, normalization, dimensionality reduction, clustering, and pattern identification to discover biologically meaningful insights
  • Implement data quality control processes and statistical analysis methods for rare-cell detection using CIARA and Monocle3
  • Develop interactive R Shiny applications for data visualization and exploratory analytics
  • Document analytical procedures and maintain data integrity, reproducibility, and workflow consistency across projects

CAMP Scholar 2025 | California Louis Stokes Alliance for Minority Participation

January 2025 - March 2025 (3 months)

  • Performed exploratory data analysis on single-cell RNA-seq datasets through data preprocessing, normalization, dimensionality reduction, clustering, and quality-control workflows
  • Developed an interactive R Shiny application for exploring and visualizing gene-expression data patterns and trends
  • Applied statistical methods for pseudotime analysis and rare-cell population identification
  • Presented data findings and communicated analytical insights to faculty and research audiences

Machine Learning Trainee | Science Coding Immersion Program (SCIP 2024), San Francisco State University

May 2024 - June 2024 (2 months)

  • Built machine-learning and data analytics workflows covering data cleaning, preprocessing, feature preparation, model training, evaluation, and interpretation
  • Applied statistical and machine-learning methods (Logistic Regression, XGBoost) to analyze biological datasets
  • Used Google Colab to develop and evaluate reproducible analytical workflows
  • Collaborated with peers to troubleshoot coding challenges and communicate analytical results

Data Specialist | CenExel ACT (Anaheim Clinical Trials)

June 2022 – August 2023 (1 year 3 months)

  • Maintained 99.8% data accuracy while processing, validating, and analyzing clinical research data
  • Managed data across multiple Electronic Data Capture (EDC) platforms and performed data quality analysis
  • Monitored quality-control checks, resolved data discrepancies, and maintained data integrity
  • Analyzed data patterns to identify inconsistencies and improve workflow efficiency
  • Collaborated with clinical research teams on data investigation and resolution
  • Developed and delivered training on data-entry procedures and EDC systems
  • Ensured data compliance and maintained confidentiality and HIPAA standards

Undergraduate Research Experience | Dr. Alison Miyamoto, California State University Fullerton

June 2021 – August 2021 (3 months)

Exploratory analysis of intracellular localization of PTBP1 (RRM2 domain)

  • Conducted exploratory data analysis of microscopy datasets investigating protein localization patterns
  • Analyzed and interpreted complex cell imaging data using ImageJ image analysis tools
  • Applied statistical analysis and data visualization to support biological discoveries
  • Presented research findings and analytical insights at a Research Symposium

Advanced ReTOOL Trainee | University of Florida

May 2020 - August 2020 (4 months)

  • Conducted literature analysis and data synthesis on the impact of intermittent fasting on breast cancer outcomes
  • Performed data collection, quality analysis, and preliminary statistical examination of research findings
  • Developed research skills in data interpretation and analytical thinking

Undergraduate Research Assistant | Museum Education: Natural History, University of Florida

January 2020 - March 2020 (3 months)

  • Conducted literature review and data analysis on museum education methodologies
  • Managed data collection efforts with quality control and preliminary analysis of educational metrics
  • Developed skills in data management and analysis using spreadsheet and database tools

STEM Learning Assistant | Miami Dade College

January 2019 - August 2019 (8 months)

  • Analyzed student performance data to identify learning improvement opportunities
  • Collaborated on academic objective development and data-driven tutoring strategies
  • Facilitated tutoring sessions using data-informed instructional approaches

📚 Education

University of California, Riverside
Bachelor of Science - Microbiology | June 2025

  • University Honors Program
  • Relevant Coursework: Genetics, Molecular Biology, Immunology, Virology, Applied Linear Algebra

Irvine Valley College
General Studies and IGETC - Transfer Certifications | June 2023

Miami Dade College
Associate of Arts – Computer Engineering | 2017 - 2019

  • Highest Honors
  • Dean's List (2017–2019)

🏆 Certifications

SnowPro Core Certification — Snowflake
Hands-On Essentials: Data Engineering Workshop — Snowflake (Building foundation for data engineering trajectory)
Hands-On Essentials: Data Warehousing Workshop — Snowflake
NASA GeneLab GL4U Intro OnDemand
Biotechnology Laboratory Assistant Skills Certificate

In Progress: Google Cloud BigQuery Fundamentals · Databricks Fundamentals (Expanding toward data engineering)


🪴 Interests

  • Open science and reproducible analytical workflows
  • Data storytelling and communicating insights from complex datasets
  • Single-cell + spatial multi-omics data analytics
  • Exploratory data analysis and pattern discovery
  • Ethical AI in biomedical research and analytics
  • Growing expertise in data engineering and ETL/ELT processes

🚀 Career Development Path

Current: Data Analyst | Data Analytics Engineer
Future Growth: Data Engineer | Analytics Engineer

I'm strategically building skills in SQL, data warehousing, and ETL/ELT processes while leveraging my strong foundation in data analytics and scientific data interpretation. My goal is to transition into data engineering roles where I can design and optimize data pipelines while maintaining the analytical rigor I bring from my research background.


📬 Let's Connect

Email: perezeconse@gmail.com
LinkedIn: linkedin.com/in/constanzaeugenia
GitHub: github.com/ceugenia


👾 Explore my repositories below! ⤵️

Pinned Loading

  1. immune-chord immune-chord Public

    An end-to-end R pipeline utilizing the BigSur algorithm for robust detection and analysis of rare cell populations in single-cell RNA sequencing data.

    R 4

  2. ptbp1-imagej-analysis ptbp1-imagej-analysis Public

    Automated ImageJ/Fiji and Python pipeline for quantifying PTBP1 and RRM2 domain subcellular localization from fluorescence microscopy images in heterokaryon assays.

    Jupyter Notebook 4