Data Engineer Job Description
A complete data engineer job description template, including where the role ends and analytics engineering begins, plus its ATS keywords.
Data Engineer is the most software-engineering-adjacent title in this group, and the closest thing to a stable definition: you move data reliably from where it's produced to where it's usable, and you're judged on distributed-systems fundamentals — schema design, failure handling, backfills, cost — more than on any specific tool. The tool list on a posting (Airflow versus Dagster, Spark versus something newer) matters much less than whether you can reason about what happens when a pipeline fails at 2 a.m.
The scope varies with something more diagnostic than company size: whether the company has a separate analytics engineer. That title split off the second half of what used to be a data engineer's job — writing and testing the SQL models that turn raw, ingested data into clean tables an analyst or data scientist can trust. Where a dedicated analytics engineer exists, "Data Engineer" narrows to ingestion, orchestration and platform reliability. Where one doesn't — still most companies under a few hundred people — the data engineer absorbs both halves, and a posting heavy on dbt and data-modeling language is telling you that up front, whatever the title says.
Because the plumbing has to work for data scientists and analysts to do anything, a data engineer's requests come from every direction and rarely arrive with a clean spec — you're building for people who often can't fully articulate what they need until they see it. That makes requirements-gathering and clear scoping as much a part of the real screening bar as any specific pipeline tool, even though postings rarely list it as a skill.
Bellamyre · Denver, CO
Full-time · Hybrid
$112,000 – $152,000
About the role
Bellamyre is hiring a Data Engineer to join the data platform team, which owns ingestion, warehousing and the pipelines everything downstream depends on. You'll build and operate the systems that move event and billing data from our production databases into the warehouse reliably, on schedule, and in a shape the analytics and data science teams can actually use.
This is a hybrid role based in our Denver office, three days a week alongside the team whose pipelines you'll be running. The platform is under active buildout, so there's more greenfield work here than at a company with a decade of pipeline debt to maintain — though there's some of that too.
What you'll do
- Design, build and operate ETL/ELT pipelines that move production data into the warehouse on a reliable schedule.
- Own pipeline failures end to end — alerting, root cause, backfill — rather than handing incidents off to whoever's awake.
- Design schemas and partitioning strategies that hold up as data volume grows, not just at today's scale.
- Work with analytics engineering and data science to understand what shape of data they actually need downstream.
- Build and maintain data quality checks that catch a broken upstream source before an analyst does.
- Manage orchestration — scheduling, dependencies, retries — across a growing number of interdependent pipelines.
- Participate in an on-call rotation covering the pipelines your team owns.
- Document data lineage well enough that someone can trace a number on a dashboard back to its source table.
What we're looking for
- Three or more years building and operating production data pipelines.
- Strong Python and SQL, including comfort optimizing a slow or expensive query.
- Hands-on experience with a workflow orchestrator such as Airflow, Dagster or similar.
- Practical experience with a cloud data warehouse such as Snowflake, BigQuery or Redshift.
- Understanding of distributed-systems fundamentals: partitioning, idempotency, backfills, failure recovery.
- Experience with version control and CI/CD applied to data pipelines, not just application code.
Nice to have
- Experience with streaming systems such as Kafka, beyond batch ETL.
- Familiarity with dbt or another transformation tool, even if analytics engineering owns it here.
- Exposure to infrastructure as code, such as Terraform, for provisioning data infrastructure.
- Experience with Spark or another distributed processing framework at meaningful data volume.
- A track record of reducing a pipeline's cost or runtime without breaking what depends on it.
Benefits
- Medical, dental and vision coverage, with employee premiums covered in full.
- 401(k) with a 4% company match, vested immediately.
- Hybrid schedule: three days a week in the Denver office.
- $2,000 annual learning budget, usable on conferences, courses or books.
- Twenty days of paid time off plus company holidays.
Salary range
As posted for this sample role. Real pay varies by employer, location and experience.
$112,000–$152,000/ yr
ATS keywords for this role
The applicant tracking system (ATS) — the recruiting software a hiring team searches and filters applicants with — will screen for these. Weight shows how central each one is to this specific posting.
Required and central (4)
Important (9)
Mentioned in passing (6)
Frequently asked questions
What's the difference between a data engineer and an analytics engineer?
Roughly, raw versus refined. A data engineer gets data into the warehouse reliably; an analytics engineer transforms it, inside the warehouse, into tested tables an analyst or dashboard can trust. At companies without a dedicated analytics engineer, the data engineer does both — heavy dbt or data-modeling language in the requirements is the tell.
Do I need a computer science degree for data engineering?
It's common but not universal. What matters more in practice is demonstrated experience with distributed systems and pipeline failure modes — a strong portfolio project or prior on-call experience often carries more weight than the degree itself, especially past the first job.
How much software engineering skill does data engineering actually require?
More than the title implies to people outside the field. Production pipelines get code-reviewed, tested and version-controlled the same as application code, and a data engineer is increasingly expected to write software that happens to move data, not scripts that happen to run on a schedule.
This was a sample. Your resume should be tailored to the real thing.
Rezi Ninja reads an actual job posting and rewrites your resume to match it, with a Ninja Score so you know it lands before you hit send. Free to start, with the AI usage included.
