About the job Remote | Data Engineer — $140,000–$180,000/year
We are sharing a full-time opportunity for an experienced Data Engineer with strong expertise in Python, SQL, Apache Spark, AWS, distributed data processing, and scalable data architecture to build and operate the infrastructure supporting AI-driven products and research initiatives.
The role will focus on designing and scaling distributed data pipelines, managing large datasets across cloud environments, and building reliable data systems that support processing, experimentation, analytics, and model development at scale.
Key Responsibilities
Scalable Data Pipeline Development
- Design, build, and maintain large-scale data pipelines
- Ingest and transform data from multiple structured and unstructured sources
- Develop reliable workflows for high-volume data processing
- Improve pipeline scalability, maintainability, and operational performance
- Support downstream analytics, experimentation, and model-development requirements
Distributed Data Processing
- Develop distributed processing workflows using Apache Spark
- Optimise large-scale data transformations and computational workloads
- Design partitioning strategies appropriate for high-volume datasets
- Identify and resolve performance bottlenecks across distributed environments
- Apply cloud-native technologies to scalable data-processing workflows
Data Architecture & Storage
- Design scalable architectures across SQL and NoSQL systems
- Build storage layers that support reliability, performance, and growth
- Evaluate database and storage technologies for different workload requirements
- Improve data accessibility while maintaining system integrity
- Apply sound architecture principles to large and complex datasets
AWS Data Infrastructure
- Design and implement cloud-native data architectures on AWS
- Support high-volume ingestion, processing, storage, and distribution
- Build infrastructure that can scale with increasing data workloads
- Improve reliability and operational efficiency across cloud data systems
- Apply AWS services appropriately across different data-processing scenarios
Python & SQL Engineering
- Write efficient Python code for data ingestion, transformation, validation, and analysis
- Develop performant SQL queries and data-processing logic
- Improve existing pipelines and transformation workflows
- Apply strong engineering standards to production data systems
- Maintain clear and reusable data-processing code
Data Quality, Monitoring & Reliability
- Ensure data quality and integrity throughout pipelines and storage layers
- Implement monitoring for data workflows and infrastructure
- Identify failures, anomalies, and processing issues
- Improve operational reliability through automation and validation
- Establish processes that support consistent and trustworthy data delivery
Cross-Functional AI & Data Collaboration
- Collaborate with AI researchers, data scientists, and engineering teams
- Support data-intensive applications and experimentation
- Provide infrastructure for model-training and evaluation workflows where relevant
- Translate research and product requirements into scalable data systems
- Adapt infrastructure as AI and data requirements evolve
Ideal Profile
- Strong professional experience in data engineering or distributed data systems
- Advanced proficiency in Python and SQL
- Hands-on experience with Apache Spark or comparable distributed-processing frameworks
- Strong experience with AWS data services and cloud-native architecture
- Experience working with both SQL and NoSQL databases
- Demonstrated experience processing and managing large-scale datasets
- Strong understanding of data partitioning and performance optimisation
- Experience designing scalable and reliable data architectures
- Familiarity with automation, orchestration, and data-pipeline monitoring
- Strong understanding of data quality and operational reliability
- Exposure to AI/ML workflows or research environments is advantageous
- Familiarity with LLM-related training, evaluation, or prompt-experimentation datasets is beneficial
- Experience with data-visualisation tools such as Matplotlib, Seaborn, or Plotly is a plus
Engagement Details
- Full-time engagement
- Fully remote
- Base compensation: $140,000–$180,000/year
- Work will involve Python, SQL, Apache Spark, AWS, distributed data processing, and scalable data architecture
- Strong experience with large-scale data systems and cloud infrastructure is central to this role
- Responsibilities will span data ingestion, transformation, storage, monitoring, and operational reliability
- The role may support AI/ML experimentation, model-development workflows, and LLM-related data infrastructure
- Data volumes, infrastructure requirements, and technical priorities may evolve as products and research initiatives scale
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy