Skip to content
View iamabhaydawar's full-sized avatar
:electron:
β΅£ Flow State Ξ¨
:electron:
β΅£ Flow State Ξ¨

Highlights

  • Pro

Block or report iamabhaydawar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
iamabhaydawar/README.md

Hello, I'm Abhay ✨ # TonyStark

I'm a Data Engineer and AI Enthusiast based in Vancouver.

I build production-grade CDC pipelines, medallion lakehouses, and streaming systems on Databricks and Snowflake.

I also create open-source AI tools β€” from intelligent LLM routers to career intelligence platforms powered by knowledge graphs.

πŸ“ My Work

What I work onπŸ›

  • Warehouse-native ELT β€” Snowflake, dbt, external stages, storage integrations
  • Lakehouse streaming β€” Delta Lake, Structured Streaming, Change Data Feed, DLT
  • Pipeline orchestration β€” Apache Airflow, Databricks Workflows, DAG design
  • Data quality β€” DLT expectations, validation gates, quarantine patterns
  • Dimensional modeling β€” SCD Type 2, fact/dim schemas, CDC upsert logic

Some cool gifs regarding my areas of interests:

Gradient Descent
(Optimization in action β€” pure Manim elegance)
Backpropagation
(How neural nets actually learn)
Gradient Descent β€” 3Blue1Brown / Manim style animation Backpropagation animation

πŸ’» Deep Dive into my Github Portfolio Projects { https://iamabhaydawar-github-io.vercel.app/ }

πŸš€ Agentic Engineering Projects

Project Description
DevRadar Β· AI Career Intelligence Platform AI-powered career intelligence platform for Indian developers. Maps personal tech stacks into a live knowledge graph, intelligently matches users to top startups, surfaces relevant hackathons, identifies skill gaps, and delivers personalized learning roadmaps + career chat. Built with React 18 (Vite), Node.js + Express, Groq (primary AI) + Claude fallback, and HydraDB for persistent memory. Features vis-network graph visualization, 4-step onboarding, 3 light themes, and graceful degradation.

Zombie CLI Β· LLM Routing Engine Multi-specialist AI routing engine that decomposes queries into subtasks and dispatches each to the best narrow model β€” Claude for code, DeepSeek R1 for math, Perplexity Sonar for research, Gemini Flash for summarization, GPT-4o for structured output, Grok-3 for fact-checking. Built on LangGraph with parallel/sequential execution, automatic fallback routing, and a cross-family verification loop with retry.

MnemOS Β· Agentic OS with Persistent Memory A visual workflow builder for desktop AI agents running in a fully containerized virtual desktop. Agents carry memory across sessions via HydraDB, recover from failures automatically, and adapt strategy based on past runs. Extends a browser-automation engine with four primitives: Remember, Recall, Recover, and Plan nodes β€” enabling graph-enhanced semantic retrieval and LLM-guided retry logic. Built for the "Agents Under Pressure" hackathon.

πŸ›  Data Engineering Projects

Project Description
UPI Transactions CDC Streaming Β· Databricks Production Change Data Capture pipeline using Delta Lake Change Data Feed. Handles INSERT/UPDATE/DELETE via a multiplier pattern (+1, -1, 0) for idempotent merchant aggregations. Hourly metrics via Delta MERGE upserts with processing_log monitoring.

HealthCare DLT Medallion Pipeline Β· Databricks DLT Patient admission analytics across Bronze β†’ Silver β†’ Gold medallion layers on Delta Live Tables. Streaming ingestion with EXPECT constraints (pk_not_null, required_fields, has_diagnosis) and ON VIOLATION DROP ROW. Three gold tables for admission trends, diagnosis prevalence, and demographics.

Ecomm Event-Driven Pipeline Β· Databricks Workflows Eight-stage pipeline triggered by file arrival across 5 source systems. Staging β†’ validation β†’ enrichment β†’ Delta MERGE. SCD Type 2 on customer dimension. Idempotent and re-runnable with automated file archival.

News Data Analysis Β· Airflow + Snowflake + GCS End-to-end pipeline: NewsAPI β†’ GCS β†’ Snowflake orchestrated by Apache Airflow. Daily DAG with pagination, Parquet landing via storage integration, and schema-inferred raw table feeding summary_news and author_activity views.

Snowflake Customer DML Medallion Β· Snowflake Medallion Architecture End-to-end Bronze β†’ Silver β†’ Gold medallion architecture on Snowflake with incremental Change Data Load, DML upserts, and schema evolution. Optimized for customer analytics using efficient merge operations, data quality checks, and production-grade incremental loading patterns.

Travel Booking SCD2 Warehouse Β· Databricks Data Warehouse Production data engineering pipeline for travel booking analytics implementing SCD Type 2 to track historical changes in customers, trips, and bookings across Bronze β†’ Silver β†’ Gold layers. Includes PyDeequ data quality validation, audit logging, Z-Order optimization, surrogate keys, and parameterized Databricks workflows.

Car Rental Batch Ingestion Β· Airflow + Snowflake Pipeline Cloud-native batch processing pipeline for car rental analytics. Orchestrated by Apache Airflow with Google Cloud Dataproc and PySpark transformations. Implements SCD Type 2 for customer dimension management and loads into a star-schema data warehouse on Snowflake for BI-ready analytics.

Pinned Loading

  1. HealthCare_DLT_Medallion_Pipeline HealthCare_DLT_Medallion_Pipeline Public

    Healthcare analytics pipeline using Databricks Delta Live Tables and a medallion architecture (bronze–silver–gold) for reliable, high‑quality clinical data transformations.

    Jupyter Notebook 2

  2. devradar devradar Public

    DevRadar - Your career. One screen. Always remembered. | WikiThon 2026

    JavaScript 1

  3. Ecomm_event_driven_dbx_Pipline Ecomm_event_driven_dbx_Pipline Public

    Event-driven data pipeline on Databricks for real-time e-commerce data processing with incremental loading, validation, enrichment, and Delta Lake operations

    Jupyter Notebook 3 1

  4. Snowflake_news_data_analysis_project Snowflake_news_data_analysis_project Public

    News data extraction and analysis pipeline using NewsAPI, Apache Airflow, GCS, and Snowflake for automated news article processing and analytics

    Python

  5. UPI_Transactions_CDC_Streaming_Analytics UPI_Transactions_CDC_Streaming_Analytics Public

    Streaming analytics pipeline for UPI transactions using change data capture to process real‑time payments data for fraud detection and behavioral insights.

    Jupyter Notebook

  6. Travel_Booking_SCD2_Warehouse_Project Travel_Booking_SCD2_Warehouse_Project Public

    Travel booking data warehouse implementing Slowly Changing Dimension Type 2 to track historical changes in customers, trips, and bookings.

    Jupyter Notebook 1