Job Openings Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)

About the job Senior Data Engineer - Databricks & Streaming - Healthcare AI (Onsite, Evening Shift, Lahore, PKR Salary)

Requirements:

  • 4+ years of experience in data engineering, with substantial production experience in Databricks.
  • Strong experience with Spark SQL, PySpark, Delta Lake, Medallion Architecture, and Delta Live Tables (DLT).
  • Hands-on experience with Structured Streaming or equivalent production-grade streaming ingestion using Azure Event Hubs, Kafka, or Kinesis.
  • Strong understanding of checkpoint recovery, watermarking, and deduplication strategies.
  • Demonstrated experience debugging source-to-warehouse data discrepancies.
  • Ability to walk through a real-world incident involving mismatched record counts and explain how the root cause was identified and resolved.
  • Proven experience integrating third-party REST APIs in production.
  • Experience handling pagination edge cases, rate and row limits, retries, and schema drift.
  • Experience with entity resolution or data matching involving messy, real-world text data.
  • Experience with a metrics or semantic layer such as Holistics AML/AQL, dbt Metrics, or LookML.
  • Working understanding of why non-additive measures cannot be reliably calculated from pre-aggregated rollups.
  • Strong SQL and Python skills, with the ability to own data pipelines end-to-end with minimal oversight.
  • Strong written and spoken English, with the ability to collaborate effectively with a US-based team asynchronously.
  • Healthcare data experience, including referrals, payer taxonomy, claims/eligibility, or other PHI-adjacent datasets.
  • Familiarity with HIPAA handling expectations.
  • Experience with voice-agent, call-center, or telephony/conversation data.
  • Familiarity with call transcripts and containment or outcome metrics.
  • Hands-on experience with Holistics, specifically AML/AQL modeling.
  • Experience with the broader Azure ecosystem beyond Event Hubs, including ADLS, ADF, and Key Vault.

Responsibilities:

  • Build and maintain resilient ingestion pipelines for third-party vendor REST APIs.
  • Work primarily with voice-AI observability and telephony platforms.
  • Handle different pagination schemes, including offset/limit and page/cursor models.
  • Manage row and rate limits through time-windowing and adaptive bisection.
  • Implement robust schema-drift handling through contract and column-presence checks.
  • Ensure alerts are triggered when fields are renamed, moved, or removed rather than silently propagating null values.
  • Own streaming ingestion from Azure Event Hubs into Databricks using Structured Streaming and/or Auto Loader.
  • Manage checkpoints and offsets, watermarking, and at-least-once deduplication.
  • Perform source-parity reconciliation across ingested and production data.
  • Investigate row counts, dropped or duplicated events, late-arriving data, and schema mismatches.
  • Identify and resolve the root cause when ingested data does not match production sources.
  • Develop and maintain Delta Lake pipelines using a Medallion Architecture (Bronze Silver Gold).
  • Use Spark SQL and PySpark to build and maintain production data pipelines.
  • Implement idempotent MERGE upserts.
  • Work with Delta Live Tables and materialized-view constraints, including CREATE OR REFRESH and LIVE references.
  • Understand and manage differences between DLT and job execution contexts.
  • Build entity-resolution pipelines for dirty, free-text data.
  • Normalize practice, provider, and payer names using regex, canonical dictionaries, fuzzy matching, confidence-scored crosswalks, and override tables.
  • Maintain the semantic and metrics layer with rigorous metric definitions.
  • Define and maintain accurate denominators, data grain, and cohort boundaries.
  • Ensure the correct handling of non-additive aggregates, including medians and percentiles that cannot be reliably supported through aggregate-aware pre-aggregation.
  • Ensure every metric remains accurate and reproducible.
  • Instrument data quality across the entire pipeline.
  • Monitor data freshness, source parity, data contracts, and other critical quality checks.
  • Build alerting mechanisms that identify data issues before they reach dashboards.