SMS · Data Platform
Initializing Data Platform…
BOOT_SEQUENCE000%
// intro.reelshreyas.mp4 · 00:00 loop

scroll to enter the data platform ↓

Shreyas Mysore Sudesh

DataEngineer.

4 years engineering high-reliability ETL/ELT — 10M+ daily records at 97%+ uptime across Azure, AWS, and GCP lakehouses.

Azure Databricks Delta Lake Snowflake · dbt Kafka Streaming
lakehouse · livev2026.06
SOURCEINGESTLAKETRANSFORMWAREHOUSEBI
streaming
10k/min
batch
100+ GB/day
warehouse
20 TB
0 yrs
Experience
0M+
Daily Records
0%
Pipeline Uptime
0%
Incidents Reduced
0+
Curated Tables
0+ GB
Daily Ingest
// 01 · about

Engineer of cloud-native data platforms.

Data Engineer with 4 years of experience engineering high-reliability ETL/ELT pipelines — reducing data incidents by 60%, sustaining 97%+ uptime across 10M+ daily records, and optimizing Spark workflows from hours to minutes at enterprise scale. Currently at BNY, previously at Tredence and Citus Infotech, with hands-on depth across Databricks, Delta Lake, Snowflake, BigQuery, Kafka, dbt, and Airflow.

Location
United States
Email
shreyasmsudesh28@gmail.com
Phone
518-956-5361
Education
M.S. BA · UAlbany · May 2026
Certifications
  • Databricks Certified Data Engineer Associate
    Earned · Active
  • Great Learning — Data Analytics Program
    Earned · Active
// 02 · stack

Skill matrix.

Production-grade tooling I reach for across ingestion, transformation, storage, and serving.

Big Data & Processing
07
Apache SparkDatabricksPySparkSpark SQLEMRUnity CatalogDelta Lake
Pipeline Orchestration & Streaming
05
Apache AirflowdbtDagsterApache KafkaSpark Streaming
Cloud Platforms
06
AWS (S3, Glue, Lambda, EMR, Redshift)Azure SynapseAzure Data FactoryADLSGCP DataflowBigQuery
Data Warehousing & Databases
06
SnowflakeAmazon RedshiftBigQueryAzure SQL DBPostgreSQLOracle
Modeling & Architecture
05
Star SchemaSCD Type 2Medallion ArchitectureCDCDelta MERGE
Programming & Scripting
08
PythonSQLPySparkBashGitGitHub ActionsDockerKubernetes
Quality, Observability & BI
07
Great Expectationsdbt testsSLA ManagementPower BITableauBigQuery MLXGBoost
// 03 · experience

Production deployments, not demos.

Three roles, one through-line: cloud-native pipelines that scale and stay reliable in production.

01

Data Engineer @ BNY

New York, NY · Jan 2026 — Current
SOURCESAIRFLOWDATAFLOWBIGQUERYDATABRICKSPOWER BI
  • Reduced data incidents 60% engineering automated quality monitoring with Python, Great Expectations, and Airflow plus dbt tests, enforced via GitHub Actions CI/CD across analytics and ML training workflows.
  • Delivered 99%+ pipeline uptime on production ETL/ELT flows processing 10M+ daily records with Python, SQL, Airflow, and GCP Dataflow, powering executive BI and downstream ML feature pipelines.
  • Engineered 50+ ML-ready features on Databricks + PySpark for fraud risk models, orchestrating with Dagster and training directly in BigQuery ML with reproducible lineage.
  • Built real-time fraud detection infrastructure on Kafka + PySpark, delivering low-latency event pipelines scoring 10M+ daily transactions for deployed anomaly detection models.
  • Engineered BigQuery semantic layers and aggregated models consumed by Power BI for fraud risk and executive KPI visibility, cutting ad-hoc engineering requests by 40%.
60%
Incidents ↓
99%+
Uptime
10M+
Daily Records
50+
ML Features
40%
Ad-hoc ↓
02

Data Engineer @ Tredence Inc.

India · Mar 2023 — Aug 2024
ORACLEADFDATABRICKSDELTA LAKEADLSPOWER BI
  • Engineered 100+ Azure Data Factory pipelines ingesting 100+ GB/day of retail, customer, transaction, and tender data from Oracle into Databricks Delta Lake, sustaining 150+ curated tables at 98% success rate.
  • Developed Bronze, Silver, and Gold Delta Lake layers using PySpark and Spark SQL, landing curated data in ADLS Silver, standardizing transformations for enterprise-wide analytics.
  • Implemented CDC-aware incremental ingestion and Delta MERGE upsert patterns in PySpark, building SCD Type 2 dimensional models to preserve historical customer and business attributes.
  • Built SQL and PySpark validation and reconciliation checks across curated tables, partnering with BI stakeholders to define data contracts — 99% on-time SLA and 30% faster resolution.
  • Delivered 20+ Power BI and Tableau dashboards on Gold-layer datasets, enabling cross-functional teams to monitor customer activity, channel performance, tender trends, and operational KPIs.
  • Containerized Spark ETL jobs with Docker and deployed via Kubernetes (AKS), enabling automated CI/CD pipelines that improved deployment consistency and reduced manual release effort.
100+
ADF Pipelines
150+
Tables
98%
Success Rate
100+ GB
Daily Ingest
20+
Dashboards
03

Data Engineer @ Citus Infotech

India · Apr 2021 — Feb 2023
SOURCESAWS S3PYSPARKSNOWFLAKEPOWER BI
  • Led the design and implementation of a centralized data lake using AWS S3 and PySpark, integrating multiple data sources with Python-based quality checks — 25% reduction in data preparation time.
  • Developed a monitoring system using Hadoop and Power BI enabling real-time visualization of pipeline metrics and alerts, reducing system downtime by 20% through early issue detection.
  • Collaborated with Data Scientists to optimize data models and enhance data accessibility, implementing security measures and AWS infrastructure optimizations that reduced operational costs by 20%.
  • Architected and deployed a scalable data warehouse on Snowflake with Python, efficiently processing 10 GB daily while maintaining optimal query performance and enabling self-service analytics for 7 business users.
25%
Prep Time ↓
20%
Downtime ↓
10 GB
Daily
7
BI Users
// 04 · projects

Featured platforms.

Open-source and personal builds that mirror enterprise patterns — Medallion, dbt, exactly-once streaming.

PROJECT · 01Terraform · BigQuery · XGBoost

NYC 311 Service Demand Data Pipeline

Terraform-provisioned BigQuery analytics pipeline processing 4.4M+ NYC 311 requests with standardized ingestion, transformation, and modeling-ready tables, plus an XGBoost predictive model achieving R² = 0.786.

TERRAFORMBIGQUERYDBTXGBOOSTLOOKER
4.4M+
Records
R² 0.786
Model Fit
Terraform
IaC
PROJECT · 02dbt · Snowflake

Snowflake Analytics Engineering Pipeline

dbt + Snowflake analytics engineering project over 5M+ subscription and attribution records — 25 analytics models and 40+ dbt tests for schema integrity, nulls, uniqueness, accepted values, and referential integrity.

SOURCESDBTSNOWFLAKETESTSANALYTICS
5M+
Records
25
dbt Models
40+
DQ Tests
PROJECT · 03Kafka · Delta Lake

Real-Time Streaming Pipeline

Kafka-to-Delta Lake streaming pipeline processing ~10,000 events/min across 3 topics with checkpointing, idempotent writes, Bronze/Silver design, and sub-5-second observed latency.

KAFKASPARK STREAMINGBRONZESILVER
10k/min
Events
< 5s
Latency
3
Topics
// 05 · architecture

Reference architectures.

Switch between the three platforms I've shipped most. Hover any node for the spec.

Azure Lakehouse
ADFDATABRICKSDELTA LAKEADLSPOWER BI
node inspector

Hover a node to inspect responsibilities.

// 06 · resume

One PDF, full archive.

latest resume

Shreyas Mysore SudeshData Engineer

Full role history, project metrics, certifications, and the complete tech stack in one PDF.

  • FORMATPDF + ATS .txt
  • UPDATED2026 · Q2
  • ROLES3 · 3+ years
  • CERTS1 active · 2 planned
  • EDUM.S. BA · UAlbany

// contact

Let's build something data-driven.

Shreyas Mysore Sudesh — Data Engineer

Shreyas Mysore Sudesh

Data Engineer

United States