Senior Data Engineer II - Data Curation
- Our Vision for AI in Pharma
- Our Current Drug Portfolio
- Our Technology & Platform
At Formation Bio, our values are the driving force behind our mission to revolutionize the pharma industry. Every team and individual at the company shares these same values, and every team and individual plays a key part in our mission to bring new treatments to patients faster and more efficiently. About the Position Formation Bio is seeking a hands-on technical leader to shape the future of our Data Curation team. This role is ideal for someone passionate about modeling, harmonizing, and unifying complex biomedical and healthcare data into high-quality, stakeholder-ready assets. You'll lead efforts to transform structured, semi-structured, and unstructured data into trusted, interoperable models that power analytics, product development, and scientific decision-making across the company. This is a high-impact role for a builder who thrives on technical depth, ontology-driven modeling, and architectural leadership. You'll set standards for how Formation Bio integrates healthcare and life sciences ontologies across diverse datasets, and you'll guide the design of the technical stack that enables interoperability, reuse, and long-term usability of curated data assets. You'll partner closely with domain experts in areas like claims, EHR, and biomedical research - ensuring their deep knowledge is translated into consistent, governed, and reusable data products. Responsibilities Technical Leadership & Strategy
- Define and communicate technical direction for the Data Curation team.
- Drive the architecture and technical stack for ontology-driven harmonization across healthcare and pharmaceutical datasets.
- Partner with domain experts (claims, EHR, pharma, research) to align technical standards across diverse datasets.
- Mentor engineers in best practices for modeling, ontology integration, and scalable curation workflows.
Data Modeling, Ontology-Driven Harmonization & Unstructured Integration
- Lead development of robust SQL/dbt models that unify complex healthcare and pharma datasets.
- Apply healthcare and biomedical ontologies (e.g., SNOMED, RxNorm, UMLS, Mondo, OMOP, FHIR) to ensure interoperability and consistent integration.
- Design scalable workflows for ontology alignment, normalization, and harmonized data product creation .
- Lead integration of unstructured data sources (clinical notes, publications, documents, scientific text) using NER, NLP, embeddings, and document parsing .
- Define patterns for linking structured and unstructured assets into a unified semantic layer that is easy to query, search, and consume.
- Establish standards for vector database usage and semantic search , ensuring embeddings and structured models are connected and interoperable.
Knowledge Integration & Architecture
- Establish architectural patterns for managing ontology mappings, ontology-driven transformations, and harmonized knowledge assets .
- Define how graph-based and relational representations complement each other for interoperability.
- Collaborate with the Data Infrastructure team to align ingestion, orchestration, and governance frameworks with curation needs.
Data Quality & Catalog Stewardship
- Enforce structural and semantic quality standards with automated checks and validation.
- Maintain and enrich the enterprise data catalog, ensuring curated datasets - structured and unstructured - are discoverable and well-documented.
- Capture and codify domain-specific knowledge into durable, governed data assets.
About You
- 7+ years of experience in data engineering, semantic modeling, or data curation, with leadership experience in technical direction.
- Proven expertise in SQL/dbt modeling and integrating healthcare and biomedical ontologies .
- Hands-on experience with ontology-driven harmonization and data model integration across heterogeneous datasets.
- Strong background in data architecture and stack design , with the ability to define standards and paved paths.
- Experience working with unstructured data : entity extraction (NER), NLP, embeddings, or document parsing.
- Familiarity with vector databases, semantic search, and knowledge graph concepts - and how to connect these with structured datasets for unified consumption.
- Comfortable with Python, orchestration tools (Dagster, Airflow), and working with diverse data types.
- Skilled at collaborating with infrastructure teams to balance semantic integration with scalable foundational tooling.
- Excited to mentor others, set high standards, and drive alignment across a multidisciplinary team.
Bonus points if you have:
- Experience with knowledge graph technologies (e.g., Neo4j, RDF/SPARQL, Cypher).
- Experience with healthcare and life sciences ontologies such as Mondo, OMOP, FHIR, SNOMED, RxNorm, UMLS .
- Experience harmonizing datasets from EHR, claims, or biomedical research domains.
- Contributions to enterprise data catalogs or metadata management frameworks.
Formation Bio is prioritizing hiring in key hubs, primarily the New York City and Boston metro areas, with additional growth in the Research Triangle (NC) and San Francisco Bay Area. Please only apply if you reside in these locations or are willing to relocate. Compensation: The target salary range for this role is: $220,000 - $280,000. Salary ranges are informed by a number of factors including geographic location. The range provided includes base salary only. In addition to base salary, we offer equity, comprehensive benefits, generous perks, hybrid flexibility, and more. If this range doesn't match your expectations, please still apply because we may have something else for you. You will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. #LI-hybrid
Recommended Jobs
Accounting Manager
Kforce's client, a very successful, well established fashion importer, seeks an Accounting Manager. Working closely with the CFO and Controller, the Accounting Manager will be involved with virtually …
Assistant to the Food & Beverage Director
Grandlife Hotels is home to two iconic downtown NYC destinations; Soho Grand Hotel and The Roxy Hotel - each known for its distinct cultural energy, renowned nightlife, and exceptional food & beverag…
Full Time Neurology Job NY
Whether you are searching for a position in your area or in another state, we have professionals to help you achieve your goals through our relationships with facilities nationwide - in rural settings…
Senior Caregiver
Looking for a male preferably Hungarian speaking gentleman to be a companion for my father living in Brooklyn. No food preparation, no house cleaning, only social interaction. Family members live in h…
Corporate Vice President - Life Valuation Reporting & Data Lead
Location Designation: Fully Remote Our New York Life culture has laid the foundation for over 175 years of commitment to our employees, agents, policy owners, and the communities where we live and…
Senior Engagement Manager, AWS Professional Services
DESCRIPTION The Amazon Web Services Professional Services (ProServe) team is seeking a highly skilled and versatile Senior Engagement Manager (EM) to join our team and lead the delivery of complex…
Software Engineer | Stablecoin
About Ramp At Ramp, we’re rethinking how modern finance teams function in the age of AI. We believe AI isn’t just the next big wave. It’s the new foundation for how business gets done. We’re investi…
Analyst - Market Access
Your Impact You will be part of the global Life Sciences Intelligence (LSi) Team as an Analyst – focused on Market Access, based in Boston, New York, New Jersey or Washington DC. McKinsey started s…
AI/ML Engineer
About Us: Savant is transforming how healthcare and life sciences organizations unlock the value trapped in unstructured medical data. Our platform combines cutting-edge large language models (LLMs)…
Staff Backend Engineer
Job Description Staff Backend / Platform Engineer Full-Time | Hybrid | New York City or Austin We’re seeking a seasoned Backend/Platform Engineer to help build the foundation of a new API-driv…