I'm Specialized in
Building and Designing Data Pipelines & Infrastructure.
SARTHAK MISTRY
DATA
MS Data Science @ Indiana University · Open to Relocation
Ingest queued Process queued Store queued Analyze queued Visualize queued Deliver queued
About Me

I'm a data engineer focused on Batch & Real-Time Pipelines, Cloud Data Infrastructure, and Data Modeling. I help teams turn raw, messy data into reliable, well-governed systems that power real decisions.

With a strong foundation in production-grade tooling and a detail-oriented approach, I build data platforms that are not just functional but scalable, observable, and maintainable — from ingestion to insight.

Sarthak Mistry
Your Photo
01 Experience
Jan 2026 – Present
Graduate Research Assistant
Indiana University, Bloomington
  • Designing a unified API layer for OpenBT to enable load/save operations across cloud storage services (Arweave).
  • Conducting validation experiments to benchmark code reliability and accuracy for the C++-to-Rust translation of OpenBT.
C++RustArweave
Aug 2025 – Present
Graduate Teaching Assistant
Indiana University, Bloomington
  • Mentored 100+ students in Management Access and Use of Big Data (INFO-I 535)
  • Designed 10+ assignments covering GCP, cloud computing, virtualization, data pipelines and modeling
Big DataGCPCloud
Jan 2025 – Sep 2025
Data Engineer
Project 990
  • Automated Pipeline for Ingesting yearly county level data from Bureau of Economic Analysis, Bureau of Labor Statistics and American Census Survey.
Python
Jul 2022 – Jul 2024
Data Engineer 1
Piramal Capital & Housing Finance, Mumbai
  • Developed batch pipelines ingesting data from Salesforce, MongoDB, SQL Server into Snowflake via Spark & Airflow
  • Improved regulatory reporting efficiency by ~40% automating a 150-feature Snowflake data mart
  • Reduced PII exposure risk with automated masking, 93% precision across 1M+ records
  • Built real-time CDC pipeline using Debezium, Kafka, and Spark Streaming
  • Automated economic data scraping using Python and Selenium, improving operational efficiency by 73%.
  • Increased ingestion speed by 46% automating AWS Lambda S3-to-Snowflake loads
SnowflakeSparkKafkaAirflowAWSDebezium
Oct 2021 – Mar 2022
Intern — Dex Analytics
NSEIT, Mumbai
  • Accelerated delivery cycles by engineering a CI/CD pipeline for dbt models in PostgreSQL, automating pre-aggregated data modeling and reducing dashboard query latency by 40%.
  • Reduced data ingestion errors by 25% by implementing automated dbt schema validations and data quality tests for candidate demographic datasets prior to Kibana visualization.
dbtPostgreSQLKibanaCI/CD
02 Projects
Architecture Architecture Diagram
E-Commerce Medallion Pipeline
GitHub
Airflow · Astronomer · DBT · Snowflake · Great Expectations
End to end data pipeline on Snowflake with a 4-layer medallion architecture (RAW → Staging → Data Vault 2.0 → Star Schema) across 17 dbt models.
  • 50+ data quality expectations across 7 suites using Great Expectations
  • Data Vault 2.0 layer with 3 hubs, 2 satellites, and 1 link table using MD5 hash-diff change detection and incremental loading.
  • 2 Airflow DAGs via Astro CLI and Cosmos, rendering each dbt model as an individual task.
  • Data Mart with RFM customer segmentation, product performance tiers, MoM growth, rolling revenue windows, and running aggregates.
Architecture Architecture Diagram
Blockchain Streaming Pipeline
GitHub
SPARK · TRINO · HUDI · KAFKA · SPARKML · BIGQUERY
Kappa architecture on a self-managed 3-node cluster for real-time Bitcoin transaction analysis.
  • KMeans clustering on BigQuery's Bitcoin dataset for anomaly scoring
  • Kafka producer ingesting live transactions via Blockchain.com websocket
  • Spark Structured Streaming applying model, stored as Hudi tables
  • Trino + Hive metastore for low-latency SQL analysis
Architecture Architecture Diagram
Data Lineage Application
GitHub
SNOWFLAKE · DJANGO · VISJS · React
Interactive visualization of table and column-level data lineage with dynamic graph traversal.
  • Complex SQL integrating Snowflake metadata views
  • Django REST APIs generating lineage graph nodes & edges
  • VisJS frontend with search, expand, hover, inspection
  • Dynamic upstream/downstream traversal logic
Scroll to explore more projects
03 Skills & Education
Python SQL Snowflake Spark Kafka Airflow dbt AWS Docker PostgreSQL JavaScript MongoDB Kubernetes SQL Server Rust Talend Streamlit ELK Stack Power BI Grafana
Expected May 2026
MS in Data Science
Indiana University
Luddy School of Informatics, Computing and Engineering
May 2022
BE in Electronics & Telecommunication
University of Mumbai
June 2022
Artificial Intelligence & Machine Learning Graduate
IBM
View Credly Badge
04 Contact