Data Engineer
Company: Todata Analytics
Location: Omaha, NE (on-site) Team: Data & Platform Engineering
About Todata
Todata Analytics is a HIPAA- and SOC 2-certified business intelligence platform serving clinical research site networks and professional services firms. Our products — including SiteGrades, CPAGrades, and Tod AI, our agentic AI analyst — are all built on a governed data layer where metrics are defined once and queried anywhere. The integrity of that layer is central to everything we deliver, and this role exists to protect and extend it.
The Role
We're looking for a Data Engineer who treats governance and foundational discipline as first-class work, not overhead. You'll own the pipelines that move client data from raw ingestion through to the conformed, permissioned datasets that power our products and our conversational AI. Because we operate in regulated environments, we hire people who are meticulous about the unglamorous parts.
What You'll Do
- Build and maintain data ingestion and transformation pipelines across our platform.
- Design and evolve the data models that power our products and analytics.
- Implement data governance and access controls appropriate to a regulated environment, keeping client data correctly isolated.
- Establish and uphold engineering fundamentals — CI/CD, consistent naming conventions, and environment separation — and hold the codebase to them.
- Partner with product and software development to translate client requirements into well-modeled, usable data.
- Support the infrastructure that makes governed data available to downstream consumers, including our AI products.
- Monitor data quality and lineage and respond to issues before they reach clients.
What We're Looking For
- 3+ years in data engineering, with hands-on Databricks experience (Spark, Delta Lake, Unity Catalog).
- Strong proficiency in SQL — comfortable writing complex queries, joins, window functions, and optimizing for performance, designing dimensional/star-schema data models from ambiguous requirements.
· Solid experience with Python for data processing (e.g., PySpark, pandas) and scripting
· Hands-on experience with Databricks (notebooks, Delta Lake, jobs, clusters)
· Excellent problem-solving and debugging skills — able to trace issues through logs, code, and data to find root causes
- Demonstrated care for data governance and multi-tenant isolation — you can speak to how you've kept datasets correctly separated and permissioned.
- Experience with CI/CD, version control, and disciplined naming/environment conventions in a data context.
- Track record of independently delivering foundational infrastructure work without close supervision.
- Clear written communication; you document what you build.
Nice to Have
- Experience in a HIPAA, SOC 2, or otherwise regulated data environment.
- Familiarity with healthcare/clinical research or financial/accounting data domains.
- Exposure to enabling AI/LLM consumers of a governed semantic layer.
- Databricks certification.
Why Todata
You'll work directly with leadership on a platform where your work is visible in the product and in client trust.
Comprehensive benefits package including medical, dental, vision, short and long-term disability, PTO, matching 401K, fitness center and golf simulator in building and more!
Core Values: Have Grit, Be Open, Be Curious, Collaborate, Make It Simple
Pay: $80,000.00 - $140,000.00 per year
Benefits:
- 401(k) matching
- Dental insurance
- Health insurance
- Life insurance
- Paid time off
- Professional development assistance
- Retirement plan
- Vision insurance
Application Question(s):
- This role is on-site at our office in Omaha, Nebraska. Are you currently located in the Omaha area or willing to relocate before the start date?
- How many years of hands-on experience do you have building production pipelines in Databricks?
Work Location: In person