Skip to content
View PratikTiwari5995's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report PratikTiwari5995

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
PratikTiwari5995/README.md

👋 Pratik Tiwari

Data Engineer | Cloud Architect | Azure • Snowflake • Databricks • Apache Airflow

LinkedIn Gmail GitHub


👨‍💻 About Me

I'm a Data Engineer specialized in designing and building production-grade, cloud-native data pipelines that transform raw data into actionable business intelligence. My expertise spans multiple cloud platforms and data orchestration frameworks, with a passion for building scalable, reliable systems that solve real-world business problems.

I excel at:

  • 🏗️ End-to-end ETL/ELT pipeline design using modern tools (Airflow, Databricks, Snowflake)
  • ☁️ Cloud-native architecture on Azure & AWS with strong data warehouse expertise
  • 📊 Data modeling & analytics using medallion architecture and dimensional modeling
  • 🔐 Data governance with Delta Lake, Unity Catalog, and security best practices
  • ⚙️ Workflow orchestration using Apache Airflow with Docker containerization
  • 📈 BI & analytics integrating Snowflake/Databricks with Power BI for business insights

🚀 Featured Projects

1. 📊 EOD Securities Pricing Analytics Platform ⭐ [LATEST]

A production-grade data engineering solution that ingests, transforms, and analyzes End-of-Day securities pricing data at scale.

Business Impact:

  • 87.5% faster ingestion - Reduced from 4 hours (manual) to 30 minutes (automated)
  • 2+ hours faster insights - Eliminated CSV bottlenecks for trading teams
  • 99.5% pipeline uptime - Automated monitoring & error handling
  • 80% error reduction - Data quality validation at every layer
  • 5,000+ daily records - Processing 26,000+ trades with 542 unique securities

Tech Stack:

  • Orchestration: Apache Airflow 2.8.0 (Docker)
  • Data Warehouse: Snowflake (4-layer: RAW → CORE → DIM → FACT → SA)
  • Cloud Storage: AWS S3 (Bronze layer)
  • APIs: Massive Stock Market API integration
  • BI & Reporting: Power BI (2 dashboards, 6 analytics views)
  • Alerting: Slack notifications on success/failure

Architecture:

Massive API  →  Airflow DAG  →  AWS S3  →  Snowflake ETL  →  Power BI Dashboards
  (5K/day)     (Download)     (Stage)    (Transform)         (Analytics)

Key Components:

  • Daily Data Pipeline: Automated stock price ingestion via Massive API
  • 4-Layer Data Warehouse: RAW (immutable) → CORE (cleansed) → DIM (dimensions) → FACT (analytics-ready)
  • 6 Subject Area Views: Market liquidity, equity performance, sector analysis, watchlist insights
  • Interactive Dashboards: Real-time trading insights, liquidity monitoring, volatility trends
  • Error Handling: Reject tables, validation checks, automated alerts

📌 GitHub Repository


2. 🛒 Ecommerce Data Pipeline (Azure) ⭐

Cloud-native ETL pipeline leveraging Azure Databricks, Delta Lake, and Power BI for real-time ecommerce analytics.

Highlights:

  • Medallion Architecture: Bronze → Silver → Gold data transformation layers
  • Unified Governance: Azure Managed Identity + Unity Catalog for secure, credentialless ADLS access
  • Data Reliability: Delta Lake time-travel & ACID transactions across all layers
  • Multi-Domain Processing: Order Items, Order Returns, Order Shipments
  • BI Integration: Power BI dashboards for real-time business metrics

Tech Stack:

  • Azure Databricks | PySpark | Delta Lake | Unity Catalog | Azure Data Lake | Power BI

📌 GitHub Repository


🛠️ Tech Stack

☁️ Cloud & Data Platforms

🔧 Data Processing & Languages

📊 Analytics & BI

🗄️ Databases & Storage

🧰 Tools & Additional Technologies


🎯 Core Competencies

Area Expertise
Data Engineering ETL/ELT pipeline design, data warehousing, real-time ingestion
Orchestration Apache Airflow, workflow automation, DAG design, error handling
Cloud Platforms Azure (Databricks, Data Lake), AWS (S3, EC2), multi-cloud strategies
Data Warehousing Snowflake, dimensional modeling, medallion architecture, query optimization
Big Data Processing PySpark, Spark SQL, distributed computing, optimization techniques
Data Governance Delta Lake, Unity Catalog, data quality, security & compliance
Analytics & BI Power BI dashboards, data storytelling, business intelligence, KPI tracking
APIs & Integrations REST API integration, data source connectors, third-party platform APIs
DevOps & Containerization Docker, container orchestration, CI/CD practices

📌 Key Achievements

  • 🏆 Built production-grade pipelines processing 5,000+ daily records with 99.5% uptime
  • 🏆 Reduced data latency by 87.5% through intelligent orchestration and optimization
  • 🏆 Designed scalable data warehouses serving 2+ executive dashboards and 6+ analytics views
  • 🏆 Implemented data governance with Unity Catalog and Delta Lake across multi-cloud platforms
  • 🏆 Created automated monitoring with Slack alerts and data quality validation
  • 🏆 Optimized Spark jobs improving query performance by 45% through best practices

📚 What I'm Currently Exploring

  • 🔍 Advanced Spark optimization techniques (partitioning, bucketing, caching strategies)
  • 🔍 Data lakehouse architecture patterns and best practices
  • 🔍 Real-time streaming with Kafka & Spark Structured Streaming
  • 🔍 Cost optimization for cloud data platforms
  • 🔍 Data observability & monitoring tools and frameworks

📬 Let's Connect

I'm always interested in discussing data engineering challenges, cloud architecture, and building systems that scale.

   


Building scalable data systems | Cloud-native architectures | Real-world impact through data

Pinned Loading

  1. Stock-ETL-Pipeline Stock-ETL-Pipeline Public

    An automated data engineering pipeline that transforms raw market data into real-time, analytics-ready insights using Airflow, AWS, and Snowflake.

    Python

  2. Ecommerce-data-pipeline Ecommerce-data-pipeline Public

    End-to-end Azure data pipeline using Medallion Architecture to process e-commerce facts (orders, returns, shipments) with Azure Databricks, Delta Lake, and Azure Data Lake Storage, delivering analy…

    Jupyter Notebook

  3. careplus-data-pipeline careplus-data-pipeline Public

    End-to-end AWS data engineering pipeline for processing support logs and ticket data using S3, Lambda, Glue, Redshift, and Athena with Power BI analytics.

    Jupyter Notebook 1

  4. Vocalin-Django-SocialMediaApp Vocalin-Django-SocialMediaApp Public

    Vocalin is a full-stack social media platform built with Django and Tailwind CSS. It allows users to register, log in securely, and share posts with full CRUD functionality. The project uses Django…

    HTML 1

  5. Lets-Chat-App- Lets-Chat-App- Public

    This is a real-time chat application built with React and Firebase. The app allows users to send and receive messages instantly, offering a seamless and interactive communication experience.

    JavaScript