BlueFlag LLP
Senior Databricks Engineer (ML or Analytics)
United States · full-time · remote
via HimalayasBachelor's degree2+ yrs
First seen Oct 2 · seen live today · via Himalayas
Skills mentioned
pythonsqlazureterraformsparkpytorchmachine learningci/cdgit
The posting, as published
BlueFlag is hiring Senior Databricks Engineers to help Department of Veterans Affairs teams move their data work to a premier data and artificial intelligence (AI) platform built on Databricks in Azure Commercial. The role is hands-on and customer-facing, much like professional services. You will work alongside VA teams to understand their on-premises workloads, rebuild them on Databricks, and teach them to use the platform with confidence.
We're hiring for two focus areas. The ML focus helps teams build, deploy, and monitor machine learning models. The Analytics focus helps teams move ETL pipelines and reporting to Databricks. Both need the same foundation: a data engineer who has built production pipelines and enjoys helping customers succeed. You'll join a friendly, collaborative team doing work that makes a real difference for Veterans and their families.
What You'll Do
Work directly with VA teams to scope use cases, plan migrations, and deliver working solutions on Databricks.
Assess on-premises workloads (SQL Server, SSIS packages, stored procedures, scheduled jobs) and adapt them to Databricks patterns in Azure Commercial.
Build and tune pipelines with PySpark, Spark SQL, Delta Lake, Lakeflow, and Azure Data Factory.
Use Unity Catalog for access, lineage, and sharing, so customer data stays governed as it moves.
Teach customers through workshops, pairing, and clear documentation so they can own their workloads.
Package reusable templates and code (Git, CI/CD, Databricks Asset Bundles) that speed up the next migration.
AI / ML focus
Build end-to-end ML workflows: feature engineering, training, evaluation, MLflow tracking, the Unity Catalog model registry, and model serving.
Help data scientists move models from local notebooks and on-premises servers into governed, monitored production.
Support generative AI use cases, such as retrieval-augmented generation (RAG) and agents, alongside predictive models.
Analytics focus
Convert SSIS and stored procedure ETL into Databricks pipelines, with reconciliation that proves the results match.
Replace on-premises reporting (SSRS reports, SSAS cubes) with Databricks SQL, AI/BI dashboards, and Power BI.
Publish reporting data from Databricks SQL warehouses, including semantic models and query performance tuning.
Why Join BlueFlag
At BlueFlag, we're passionate about leveraging cutting-edge technology to make a real difference. You'll be at the forefront of cloud innovation, working on projects that directly impact people's lives. We offer a high-growth, entrepreneurial environment that values fresh ideas and authentic teamwork.
If you're ready to take your data engineering career to new heights and contribute to meaningful projects that push the boundaries of technology, we want to hear from you. Join BlueFlag and be part of a team that's shaping the future of cloud solutions!
Requirements
7+ years in data engineering, including 2+ years building production workloads on Databricks
Strong proficiency in Python, SQL, PySpark, and Spark SQL
Hands-on experience with Delta Lake, lakehouse (medallion) design, and Spark performance tuning
Hands-on Unity Catalog experience: catalogs, grants, lineage, and row- and column-level security
Experience migrating on-premises workloads (e.g., SQL Server, SSIS, stored procedures) to the cloud
Experience with Azure services, including ADLS Gen2, Azure Data Factory, Entra ID, and Key Vault
Experience with Git, CI/CD, and infrastructure as code (e.g., Databricks Asset Bundles, Terraform)
Customer-facing consulting or professional services experience, including requirements sessions, workshops, and training
Clear written and verbal communication with technical and non-technical audiences
Bachelor's degree in computer science, information systems, or a related field.
US Citizen: Must be a citizen of the United States
Security Clearance: Must be able to obtain a public trust clearance. Must be eligible to work in the United States.
AI / ML focus
Experience training, evaluating, and deploying models with MLflow and common frameworks (e.g., scikit-learn, XGBoost, PyTorch)
Analytics focus
Experience with Databricks SQL and Power BI data modeling (star schemas, semantic models, DAX)
Desired
Databricks certifications (Data Engineer Professional; Machine Learning Professional or Data Analyst Associate by focus)
Experience with healthcare data (HL7, FHIR, OMOP) and HIPAA requirements for PHI
Experience in Azure Government or other FedRAMP High environments
Experience upgrading Hive metastore workloads to Unity Catalog (e.g., UCX)
AI / ML focus : Mosaic AI Vector Search, Agent Framework, or model monitoring
Analytics focus : Microsoft Fabric, Pyramid Analytics, or Genie
Benefits
Competitive salary
Generous annual leave and paid holidays
Comprehensive group health and dental plans
401(k) with company match
Life insurance and AD&D coverage
Ongoing training and professional development opportunities
Originally posted on Himalayas