Job Detail

Job Position - Department

Data Engineer - Information Technology

Experience

1–4 years in data engineering, backend data pipelines, or closely related software/data roles

Education

Bachelorʼs degree in Computer Science

Last Date

30-Aug-2026

Job Description

Job summary.

We are building an on-premises analytical data platform to bring together clinical and operational data from multiple hospital sites into a governed lakehouse (raw, curated, and serving layers). This role focuses on pipelines, data reliability, and platform operations — turning captured data into trusted, linked, analysis-ready datasets.

You will work as part of a data team under established architecture and quality standards. A platform steering committee and senior consultant provide design direction, phase planning, and review; you will implement, document, and operate the ingestion and transformation layer with growing independence over time.

Key responsibilities

Data ingestion & platform pipelines

Design, build, and maintain batch and change-capture ingestion from operational databases across multiple sites. Implement orchestrated workflows for scheduled extraction, landing, transformation, and promotion between lakehouse layers.

Ensure pipelines are idempotent, observable, and recoverable (retries, checkpoints, replay where appropriate).

Work with database administrators on read-only access, change-log readiness, and safe extract windows.

Lakehouse layers (raw → curated)

Manage immutable raw landing and curated (silver) datasets following medallion-style layering.

Implement validation, cleansing, typing, and conformed models using transformation-as code practices.

Support slowly changing history for key demographic and reference entities where attributes change over time.

Apply configuration-driven pipeline definitions (declarative specs reviewed in version control) rather than one-off scripts per table.

Cross-site identity & data linking Implement logic to unify records across sites (e.g. patients, providers, facilities) using defined matching rules and steward review for ambiguous cases.

Maintain bridge and reference structures that map source identifiers to enterprise identifiers with full audit trail.

Data quality, reconciliation & operations

Run and automate reconciliation between source systems and platform copies (counts, keys, samples, freshness).

Respond to pipeline failures and data drift using runbooks and escalation paths.

Contribute to schema change handling (additive vs breaking changes) with documentation and alerts.

Monitor service levels (lag, success rate, data freshness) via operational dashboards.

Documentation & collaboration

Maintain technical runbooks, pipeline documentation, and change records on the same day as changes.

Partner with the Analytics Engineer on catalog entries, data dictionary fields, and lineage metadata.

Support clinical and operational data stewards with technical fixes; business decisions on duplicates and definitions stay with stewards.