via Teamtailor
$90K - 140K a year
Build and operate scalable data ingestion, ETL/ELT, and orchestration pipelines ensuring data quality and real-time processing.
5+ years data engineering experience with Python, SQL, PySpark, data modeling, lakehouse architectures, and real-time data processing.
This is a hands-on building role: you turn raw, messy fabrication data into the clean, well-modeled, AI-ready datasets that our AI/ML and analytics workloads run on 🚀 🧑🏻💻 Responsibilities: Pipeline Development & Operation: Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines (batch and real-time streaming) within Microsoft Fabric and cloud lakehouse environments. Real-Time Data Ingestion: Design and implement low-latency, real-time data ingestion flows to support live operational analytics and streaming workloads. Data Modeling & Layering: Implement layered (medallion-style: Bronze/Silver/Gold) architectures using PySpark/SQL with idempotent, backfillable, and incrementally loaded jobs. Data Quality & Governance: Apply deduplication, normalization, schema validation, and lineage tracking to ensure downstream data is high-quality, trustworthy, and audit-ready. AI & Analytics Readiness: Deliver feature-ready, curated datasets to support business intelligence, analytics, vector search, and AI/ML agentic workloads. Observability & Reliability: Establish testing, monitoring, and pipeline observability (freshness, volume, schema drift) with clear alerting to resolve failures proactively. Tooling & AI Development: Utilize AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for writing pipelines, query tuning, and data transformation scripts. 🤝 If you have: Experience: 5+ years of hands-on data engineering experience building and operating production data pipelines at scale. Core Technical Stack: Strong proficiency in Python, SQL, and PySpark / Apache Spark, backed by solid software engineering fundamentals (Git, CI/CD, unit/integration testing). Real-Time Data Processing: Demonstrated hands-on experience implementing real-time data ingestion and streaming pipelines (not limited to batch processing). Data Architecture & Modeling: Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture). Platform Experience: Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred. Data Quality & Observability: Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement. AI Tooling: Active experience leveraging AI-assisted development tools (Cursor, Copilot, Claude) to accelerate engineering velocity. 🦾 It’s a plus: Hands-on experience with Microsoft Fabric (Fabric Lakehouse, Data Factory, Synapse Analytics). Experience extracting data from document stores / NoSQL databases (specifically MongoDB / MongoDB Atlas and Change Streams / CDC). Streaming frameworks experience (Event Hubs, Kafka, Spark Structured Streaming). Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads. AEC / Construction / MEP domain experience. This call is made within the framework of Law 19.691 on the Promotion of Employment for Persons with Disabilities, including individuals registered in the National Registry of Persons with Disabilities of the Ministry of Social Development
This job posting was last updated on 9/28/2026