Hi There! I'm Mani Kumar Singh. A Data Engineer Building Scalable Pipelines & Lakehouse Systems

Available for hire & consulting
Snowflake SnowPro Core dbt Certified Developer GCP Professional Data Engineer Medallion Architecture Fortune 500 Delivery
3+

Years of Experience

Senior Data Engineer specializing in scalable lakehouse architectures, production Snowflake & dbt pipelines, and high-reliability data systems for Fortune 500 enterprises.

Lakehouse & Warehousing

Architecting scalable cloud data warehouses and lakehouses on Snowflake and GCP BigQuery with optimized partitioning, clustering, and dimensional models.

Automated ETL/ELT & dbt

Building robust modular data transformation pipelines using dbt, SQL, and PySpark following the Medallion Architecture (Bronze → Silver → Gold).

Governance & Automation

Implementing automated CI/CD deployment workflows, schema validation, and stringent data quality checks that eliminate 20+ hours of weekly manual toil.

Featured Engineering Work

Production Data Systems

Enterprise Solutions • High Reliability
Swipe to explore projects
Fortune 500 Client 300+ Queries → 105 dbt Marts

Customer Reward Analytics & BI Rationalization

Reverse-engineered 300+ complex SQL queries into 105 centralized dbt mart models on Snowflake, significantly reducing cloud compute costs while ensuring data integrity.

Snowflake dbt Core SQL Sigma BI Python
Retail Enterprise Real-Time Alerts & DQ

Enterprise E-Commerce & Retail Delivery Platform

Architected end-to-end retail data pipelines with automated dbt tests and native Snowflake Alerts, establishing an INFORMATION_SCHEMA DQ system.

Snowflake dbt Core SQL Snowflake Alerts GitLab CI/CD
Cloud Lakehouse Bronze → Silver → Gold

GCP Cloud-Based Medallion Retail Lakehouse

Cleaned and transformed multi-layer retail transactional datasets across Medallion architecture using Dataproc, Dataprep, and BigQuery on Google Cloud Platform.

Google Cloud (GCP) BigQuery Dataproc Dataprep dbt Python
Data Lakehouse Design

End-to-End Medallion Pipeline Flow

Raw Ingestion → Conformed → Curated Business KPIs
1. Ingestion Layer 2. Bronze (Raw Lake) 3. Silver (Cleansed) 4. Gold (Analytics Store)
Swipe pipeline stages (1 → 4)
Step 1: Sources

Multi-Source Ingestion

Batch & streaming ingestion from retail POS logs, marketing APIs, transactional databases, and event streams.

REST APIs GCP Pub/Sub Cloud Storage
Step 2: Bronze Layer

Raw Data Lake

Immutable append-only raw storage. Preserves source fidelity, lineage audit history, timestamps, and schema-on-read flexibility.

Google Cloud Storage Snowflake Staging Parquet
Step 3: Silver Layer

Cleansed & Conformed

Deduplication, schema enforcement, data type casting, business logic transformations, and data quality testing via dbt models.

dbt Core Snowflake PySpark / Python SQL
Step 4: Gold Layer

Curated Business Marts

Star & snowflake dimensional schemas optimized for low-latency BI queries, executive reporting, and downstream ML features.

Snowflake Marts GCP BigQuery Sigma BI Tableau
Career Progression

Professional Experience

3+ Years Production Data Engineering
Tredence Inc. Logo

Tredence Inc.

Promoted to Consultant
Data Engineering • Full-time
Sept 2023 – Present • 3+ yrs
Consultant
July 2026 – Present Current Role Bengaluru, India
Customer Reward & Retail Analytics Platform (Fortune 500 Client)
Tech Stack: Snowflake, dbt Core, SQL, Python, GitLab CI/CD, Confluence
  • End-to-End E-Commerce & Retail Delivery: Architected production data pipelines for omnichannel retail domains (online/offline store transactions, e-commerce orders, customer demographics, and store analytics). Managed full data lifecycle from source extraction and profiling to dimensional modeling and delivering governed business tables.
  • BI Rationalization & Query Consolidation: Reverse-engineered 300+ redundant, complex custom SQL queries powering Sigma BI reward reports, consolidating them into 105 centralized, highly optimized dbt mart models, significantly cutting cloud compute expenditure while guaranteeing data integrity.
  • Snowflake Alert & Monitoring Framework: Developed an automated real-time monitoring system combining dbt tests with native Snowflake Alerts to intercept source discrepancies before loading into target tables, instantly notifying stakeholders.
  • Automated Data Quality (DQ) System: Leveraged Snowflake INFORMATION_SCHEMA and scheduled audit tasks to enforce schema validation, referential integrity, and business rule conformance.
  • Production Support & Incident RCA: Owned pipeline uptime SLAs, conducted root cause analysis (RCA) on production incidents, implemented automated archive models for missing transactions, and mentored junior engineers on dbt standards.
Analyst
Sept 2023 – June 2026 2 yrs 10 mos Bengaluru, India
  • Enterprise E-Commerce & Retail Pipeline Engineering (Fortune 500 Client):
    Tech Stack: Snowflake, dbt Core, SQL, Python, GitLab CI/CD
    Built and maintained core production data pipelines for enterprise retail transactions, inventory updates, and store-level feeds. Developed dbt transformation models, executed regular data reconciliation routines, and provided production support and incident resolution.
  • Marketing Campaign Analytics Platform (Fortune 500 Client):
    Tech Stack: SQL, Power BI, SQL Server, Python
    Engineered automated ingestion and transformation pipelines for omnichannel marketing campaign data (push notifications and email events). Built executive Power BI dashboards tracking sales, YoY revenue growth, and campaign performance, discovering high-value customer patterns that drove a 20%+ marketing-to-sales conversion uplift.
  • T-Discover — Enterprise AI Product Development:
    Tech Stack: Azure Databricks, Azure Data Factory (ADF), PySpark, Python, ML Similarity Models
    Core engineer building Tredence's flagship AI product to revolutionize BI report rationalization and eliminate redundant data assets using ML similarity scoring. Optimized Databricks Spark clusters, tuned PySpark functions, and automated production workflows with ADF.
  • GCP Cloud-Based Medallion Retail Lakehouse:
    Tech Stack: Google Cloud (GCP), BigQuery, Dataproc, Dataprep, Cloud Storage, SQL, Python
    Cleaned and transformed multi-layer retail transactional datasets across Medallion architecture (Bronze → Silver → Gold) using Dataproc, Dataprep, and BigQuery, establishing automated execution logging and technical documentation.

Bijli Solutions

Data Engineering & Analytics • Internship
June 2023 – July 2023 • 2 mos
Data Engineering & Analytics Intern
June 2023 – July 2023 2 mos Remote / India

Built automated data extraction and transformation pipelines using Python and SQL. Streamlined reporting and audit reconciliation routines, reducing analysis turnaround time by 60% and eliminating 20+ weekly manual operational hours.

Academic Foundation

NIT Jalandhar

Bachelor of Technology (B.Tech)

Dr. B. R. Ambedkar National Institute of Technology Jalandhar

Graduated 2023 4 Years Full-time
Institute of National Importance

Premier technical institute established under the NIT Act by the Government of India, recognized for engineering and research excellence.

Core Engineering Foundation
  • Computational Logic: Algorithms, data structures, and distributed computation.
  • Data Foundations: Relational Database Management Systems (RDBMS) & SQL.
  • Quantitative Analysis: Engineering mathematics, statistics, and probability.
  • Analytical Problem Solving: Structured systems design & architecture principles applied directly to enterprise data engineering.
Enterprise Experience

Delivering for Fortune 500 Retailers

Architecting and operating high-throughput production data pipelines on Google Cloud and Snowflake, ensuring critical sales, customer, and marketing datasets are delivered with high availability.

20+ Weekly Manual Hours eliminated via automated CI/CD workflows
60% Reduction in data audit and reporting turnaround times
Full Lifecycle Delivery from bronze ingestion to gold reporting marts
Peer Endorsements

“Mani is very supportive and always ready to help. He never panics, handling complex pipeline challenges with calm while sharing valuable technical lessons that enrich team discussions.”

MK
Manish Kumar
Colleague & Engineering Collaborator

“Mani is a highly dependable and knowledgeable engineer. We worked closely together on enterprise data delivery, dbt modeling, and automated pipeline governance at Tredence.”

SS
Satyam Singh
Data Engineer, Tredence Inc.
Technical Competencies

Core Tech Stack & Tools

Enterprise Tools • Zero Fluff
Swipe to explore competencies
Cloud & Warehousing
Google Cloud Snowflake Databricks Azure BigQuery
ETL & Modeling
dbt Core Medallion Architecture ETL / ELT Dimensional Modeling Data Warehousing
Code & Automation
Python SQL (Advanced) PySpark CI / CD Git & DevOps
Quality & Analytics
Governance Sigma BI Power BI Quality Auditing
Verified Credentials

Certifications & Honors

Industry Recognized • Enterprise Excellence
Snowflake Logo
Snowflake

SnowPro Core Certified

Certified in cloud data warehousing, performance optimization, data sharing, security governance, and scalable storage architecture.

Verify Credential
dbt Labs Logo
dbt Labs

dbt Certified Developer

Demonstrated expertise in building modular, tested, and documented analytics engineering transformation workflows with dbt Core.

Verify Credential
Google Cloud Logo
Google Cloud

Professional Data Engineer

Certified in architecting data processing systems, operationalizing machine learning models, and building scalable BigQuery & GCP pipelines.

Verified Professional
Award Badge
Tredence Inc. • June 2025

PAT ON THE BACK Award

Honored with enterprise recognition for outstanding engineering performance, dedication, and exceptional contribution to retail client projects.

Excellence Honor

Snowflake & dbt

Medallion Architecture

Databricks Lakehouse

Google Cloud BigQuery