This is a fully hands-on, individual contributor Data Engineering role embedded within a data infrastructure initiative at a large enterprise. You will own the full lifecycle of AWS-based data pipelines, from ingesting and streaming source data through to delivering clean, consumer-ready datasets for downstream use.
Stream and process source data from mainframe and legacy systems, loading it into AWS S3.
Design and implement ETL/ELT workflows using AWS Glue.
Perform data reconciliation, validation, and quality checks to ensure data accuracy and reliability.
Curate and transform data, provisioning clean datasets through AWS Aurora and RDS (PostgreSQL).
Manage S3 storage including data retention policies, archival strategies, and lifecycle management.
Automate workflows and manage deployments using GitHub and GitHub Actions-based CI/CD pipelines.
Leverage AI tools to improve engineering productivity and automate data workflows.
5 or more years of professional Data Engineering experience building and delivering data pipelines, ETL/ELT workflows, or data platform solutions.
Hands-on production experience with AWS Glue for designing and implementing ETL/ELT workflows.
Strong working knowledge of AWS cloud services including S3, RDS/Aurora, and Glue.
Demonstrated experience building and maintaining data streaming pipelines using Kafka.
Experience implementing data reconciliation, data quality checks, and validation processes.
Proficiency with GitHub repository management and GitHub Actions for CI/CD automation.
Ability to take ownership of deliverables and drive work to completion with minimal oversight.
Familiarity with MongoDB or other NoSQL databases is a plus.
Compensation & Benefits
This role pays
$65/hr on a W2 basis
.
Location
This role is
100% remote
.
Originally posted on
Himalayas
Browse by category