I'm a Data Engineer and AI Enthusiast based in Vancouver.
I build production-grade CDC pipelines, medallion lakehouses, and streaming systems on Databricks and Snowflake.
I also create open-source AI tools β from intelligent LLM routers to career intelligence platforms powered by knowledge graphs.
What I work onπ
- Warehouse-native ELT β Snowflake, dbt, external stages, storage integrations
- Lakehouse streaming β Delta Lake, Structured Streaming, Change Data Feed, DLT
- Pipeline orchestration β Apache Airflow, Databricks Workflows, DAG design
- Data quality β DLT expectations, validation gates, quarantine patterns
- Dimensional modeling β SCD Type 2, fact/dim schemas, CDC upsert logic
Some cool gifs regarding my areas of interests:
| Gradient Descent (Optimization in action β pure Manim elegance) |
Backpropagation (How neural nets actually learn) |
|---|---|
![]() |
![]() |
π» Deep Dive into my Github Portfolio Projects { https://iamabhaydawar-github-io.vercel.app/ }
| Project | Description |
|---|---|
| DevRadar Β· AI Career Intelligence Platform | AI-powered career intelligence platform for Indian developers. Maps personal tech stacks into a live knowledge graph, intelligently matches users to top startups, surfaces relevant hackathons, identifies skill gaps, and delivers personalized learning roadmaps + career chat. Built with React 18 (Vite), Node.js + Express, Groq (primary AI) + Claude fallback, and HydraDB for persistent memory. Features vis-network graph visualization, 4-step onboarding, 3 light themes, and graceful degradation. |
| Zombie CLI Β· LLM Routing Engine | Multi-specialist AI routing engine that decomposes queries into subtasks and dispatches each to the best narrow model β Claude for code, DeepSeek R1 for math, Perplexity Sonar for research, Gemini Flash for summarization, GPT-4o for structured output, Grok-3 for fact-checking. Built on LangGraph with parallel/sequential execution, automatic fallback routing, and a cross-family verification loop with retry. |
| MnemOS Β· Agentic OS with Persistent Memory | A visual workflow builder for desktop AI agents running in a fully containerized virtual desktop. Agents carry memory across sessions via HydraDB, recover from failures automatically, and adapt strategy based on past runs. Extends a browser-automation engine with four primitives: Remember, Recall, Recover, and Plan nodes β enabling graph-enhanced semantic retrieval and LLM-guided retry logic. Built for the "Agents Under Pressure" hackathon. |
| Project | Description |
|---|---|
| UPI Transactions CDC Streaming Β· Databricks | Production Change Data Capture pipeline using Delta Lake Change Data Feed. Handles INSERT/UPDATE/DELETE via a multiplier pattern (+1, -1, 0) for idempotent merchant aggregations. Hourly metrics via Delta MERGE upserts with processing_log monitoring. |
| HealthCare DLT Medallion Pipeline Β· Databricks DLT | Patient admission analytics across Bronze β Silver β Gold medallion layers on Delta Live Tables. Streaming ingestion with EXPECT constraints (pk_not_null, required_fields, has_diagnosis) and ON VIOLATION DROP ROW. Three gold tables for admission trends, diagnosis prevalence, and demographics. |
| Ecomm Event-Driven Pipeline Β· Databricks Workflows | Eight-stage pipeline triggered by file arrival across 5 source systems. Staging β validation β enrichment β Delta MERGE. SCD Type 2 on customer dimension. Idempotent and re-runnable with automated file archival. |
| News Data Analysis Β· Airflow + Snowflake + GCS | End-to-end pipeline: NewsAPI β GCS β Snowflake orchestrated by Apache Airflow. Daily DAG with pagination, Parquet landing via storage integration, and schema-inferred raw table feeding summary_news and author_activity views. |
| Snowflake Customer DML Medallion Β· Snowflake Medallion Architecture | End-to-end Bronze β Silver β Gold medallion architecture on Snowflake with incremental Change Data Load, DML upserts, and schema evolution. Optimized for customer analytics using efficient merge operations, data quality checks, and production-grade incremental loading patterns. |
| Travel Booking SCD2 Warehouse Β· Databricks Data Warehouse | Production data engineering pipeline for travel booking analytics implementing SCD Type 2 to track historical changes in customers, trips, and bookings across Bronze β Silver β Gold layers. Includes PyDeequ data quality validation, audit logging, Z-Order optimization, surrogate keys, and parameterized Databricks workflows. |
| Car Rental Batch Ingestion Β· Airflow + Snowflake Pipeline | Cloud-native batch processing pipeline for car rental analytics. Orchestrated by Apache Airflow with Google Cloud Dataproc and PySpark transformations. Implements SCD Type 2 for customer dimension management and loads into a star-schema data warehouse on Snowflake for BI-ready analytics. |




