Ask most teams what a data engineer does and you get a list of tools: SQL, Airflow, dbt, Spark, a warehouse. Those are real, and SQL remains one of the most-used languages anywhere, sitting at 58.6% in the developer survey (Stack Overflow 2025). But tools are not the job. The job is trust. A data engineer's real product is a number that a decision-maker can rely on, and the hard part is not building the pipeline. It is knowing when the pipeline is lying.
Key Takeaways
- A data engineer's real deliverable is trustworthy data, not pipelines for their own sake.
- The senior skill is data quality and failure detection, not tool count.
- Test how they catch a pipeline that runs successfully but produces wrong numbers.
- SQL fluency is table stakes; reasoning about correctness is the differentiator.
The Failure Mode That Defines the Role
Software fails loudly. A data pipeline fails quietly. It runs to completion, reports success, and delivers a dashboard where revenue is understated by 12% because an upstream schema changed and a join silently dropped rows. Nobody gets paged. The CEO makes a decision on a wrong number, and you find out weeks later. That silent-wrong-data scenario is the thing a good data engineer spends their energy preventing, and it is exactly what a tool-focused interview never tests.
So build your interview around correctness under quiet failure. A candidate who only talks about moving data fast, without ever mentioning validation, freshness checks, or reconciliation, is telling you they have not been burned yet. The ones worth hiring have war stories about a number that was wrong and how they now stop it from happening.
The Questions That Matter
Get concrete about trust. Ask how they would know if a pipeline that ran successfully actually produced correct output. Ask what checks they put between raw data and the dashboard. Ask what they do when a business user says "this number looks wrong."
| Weak signal | Strong signal |
|---|---|
| Lists tools and frameworks | Talks about data quality checks |
| "The pipeline ran, so it's fine" | "Success does not mean correct" |
| Optimizes for speed only | Reconciles against a source of truth |
| No testing story | Tests transformations like code |
A Concrete Screen
Describe a real situation and let them investigate. Daily revenue in a dashboard dropped 30% overnight. Engineering swears nothing changed. Walk me through it. A strong data engineer does not guess. They check whether the drop is real or an artifact: did row counts change, did an upstream source change its schema, is a timezone boundary double-counting or dropping a day, did a join start losing records. A weaker candidate speculates about the business ("maybe sales fell") before checking whether the number itself can be trusted. The instinct to distrust the pipeline first is the whole job.
The Honest Counterpoint
A very early startup may not need a dedicated data engineer at all. When your data fits in one database and your "pipeline" is a nightly query, a backend engineer or a technical analyst can carry it, and hiring a specialist gives you someone building elaborate infrastructure for data that does not yet justify it. The signal that you need the role: multiple sources that have to be joined, dashboards that leadership actually trusts and acts on, or data quality problems that keep embarrassing you. Before that, resist the urge to over-build.
Cost and Sourcing
A senior data engineer in the US commonly runs $150 an hour or more, and demand has climbed with the AI wave, since clean data is the input to every model. Nearshore in Latin America, senior data engineers land roughly $60 to $95 an hour, with overlap that matters because data issues are often urgent and cross-functional (nearshore staff augmentation for CTOs). Whoever you hire, screen for the instinct to distrust a green pipeline, because that habit is what separates a data engineer who ships dashboards from one who ships dashboards you can bet the quarter on. See available engineers.
Frequently Asked Questions
What does a data engineer actually do?
Builds and maintains the systems that move and transform data, and, more importantly, ensures the resulting numbers are trustworthy. The core skill is catching quiet failures where a pipeline runs but produces wrong data.
How do I test a data engineer candidate?
Give them a scenario where a metric changed unexpectedly and watch whether they investigate the data's correctness before speculating about the business. Distrusting a successful pipeline is the key signal.
When does a startup need a dedicated data engineer?
When you have multiple data sources to join, dashboards leadership relies on for decisions, or recurring data quality problems. Before that, a backend engineer or analyst often suffices.
How much does a senior data engineer cost?
In the US, commonly $150 an hour or more. Nearshore in Latin America, roughly $60 to $95 an hour at the same seniority.
The Bottom Line
A data engineer is not a person who knows Airflow. They are the person who makes sure the number on the dashboard is one you can act on, and who assumes a successful pipeline might still be wrong until proven otherwise. Test for that reflex, hire it when your data actually justifies the role, and you will trust your own numbers a great deal more.
Roberto Espinoza is CEO of Ruzora, which helps US startups hire pre-vetted senior LATAM engineers, with a vetted shortlist in 72 hours. See available engineers.
