I am a Data Platform Engineer and MLOps Architect specializing in production-grade distributed lakehouses, high-throughput streaming pipelines, and scalable machine learning infrastructure.
- ποΈ Core Architectures: Medallion Lakehouses (Delta Lake, Databricks), Real-Time Streaming (Kafka, PySpark), and Workflow Orchestration (Apache Airflow).
- π€ AI & Inference: Tabular Foundation Models, Recommendation Ensembles (SASRec, LightGCN), Vector Search (
pgvector), and Local LLM RAG pipelines. - π₯ Industry Domains: Clinical Health Informatics (FHIR R4, OMOP CDM v5.4, ABDM compliance), Capital Markets, and Media Recommendation Platforms.
π₯ AI-Healthcare-System
Enterprise AI Healthcare Lakehouse & Clinical Intelligence Platform
- Architecture: PySpark Medallion Lakehouse, Apache Airflow pipelines, FHIR R4 / OMOP CDM v5.4 compliance, and HIPAA-ready FastAPI backend.
- ML & RAG: TabICLv2 Tabular Foundation Models, calibrated CatBoost/XGBoost ensembles with 95% Conformal Prediction sets, 10-year multi-organ digital twin simulator, and local Ollama LangGraph multi-agent RAG.
- Live Links: π€ Hugging Face Live Interactive Space Β· π€ Model Hub (16 Weights) Β· β GitHub Repository
Real-Time AI Media Recommendation Engine & Unified Data Intelligence Platform
- Scale: Ingests and processes 21M+ real records (1M+ TMDB movies and 20M+ MovieLens ratings) across a Databricks Serverless Medallion Lakehouse.
- Deep Learning & Search: SASRec sequential transformers, LightGCN graph embeddings, and a 10-shard Neon Serverless
pgvectorHNSW cluster (<5ms query latency). - Live Links: π Live Cinema Portal Demo Β· π€ Hugging Face Space UI Β· β GitHub Repository
π Career Keywords & Technical Index (Search Engine Optimization)
Core Specializations: Lead Data Engineer, Senior Data Architect, Big Data Engineer Portfolio, MLOps Engineer, Lakehouse Architect, Python Developer, PySpark Specialist, Distributed Systems Engineer. Distributed Computing & Lakehouse: Apache Spark, PySpark Streaming, Delta Lake, Apache Iceberg, Apache Airflow, Databricks Medallion Lakehouse, Data Vault 2.0, SCD Type 2, Liquid Clustering, Z-Order Optimization, Great Expectations. Machine Learning & Search: Recommendation Engines, SASRec, LightGCN, Graph Neural Networks, PyTorch, pgvector HNSW Indexing, Vector Databases, Conformal Prediction, Hugging Face Spaces, Ollama Local Inference, LangGraph Multi-Agent RAG. Cloud & Infrastructure: Amazon Web Services (AWS EMR, S3, Glue, Athena, RDS, ECS), Docker Containerization, Kubernetes Orchestration, CI/CD GitHub Actions, PostgreSQL, Redis Streaming.






