IQVIA
IQVIA
IQVIA combines healthcare data analytics, advanced technology, and consulting to help organizations address complex challenges across the life sciences and healthcare sectors. Its Connected Intelligence platform brings together data and artificial intelligence to support clinical research, technology development, and the advancement of new therapies. The company’s teams work across analytics, consulting, and healthcare technology, offering opportunities connected to research, digital innovation, and efforts to improve patient outcomes.

Data Engineer, COA Accelerator — IQVIA Spain Remote

Build AI-ready clinical outcomes data for IQVIA’s healthcare intelligence business. Develop pipelines, retrieval systems, and governed knowledge assets supporting COA strategy.

Description

  • Build and maintain data infrastructure for IQVIA’s AI-enabled COA strategy and COA Accelerator.
  • Manage ingestion, transformation, normalization, enrichment, indexing, versioning, and governance across public, proprietary, and client data.
  • Create pipelines for structured and unstructured sources, including documents, databases, APIs, clinical trial registries, regulatory materials, publications, and internal repositories.
  • Convert source content into standardized, searchable, AI-ready formats for evidence retrieval, citation, recommendations, and expert review.
  • Develop document-processing workflows covering parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality assurance.
  • Support the knowledge layer with COA metadata, psychometric evidence, therapeutic-area mappings, endpoint usage, regulatory precedent, and scientific evidence.
  • Integrate public evidence from trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, and qualification documents.
  • Ingest proprietary knowledge, publications, thought leadership, and expert-authored content.
  • Prepare datasets for retrieval-augmented generation through chunking, embeddings, indexes, metadata filters, and source-reference structures.
  • Partner with AI engineers to improve retrieval precision, recall, relevance, and citation accuracy.
  • Implement hybrid retrieval across semantic and keyword search, structured database queries, and metadata filtering.
  • Preserve traceability between AI-generated results and their source documents.
  • Establish data-quality controls and audit trails for ingestion, transformation, updates, deletions, permissions, and downstream use.
  • Coordinate with legal, security, compliance, product, and domain stakeholders on licensing, privacy, intellectual property, contractual, and governance matters.
  • Work cross-functionally with AI Engineering, Product, COA Science, and Software Engineering.

Requirements

  • A degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical discipline.
  • Experience designing, developing, and maintaining pipelines for structured and unstructured data.
  • Strong proficiency in Python and SQL.
  • Experience with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL or ELT workflows.
  • Background working with large volumes of text-rich documents.
  • Solid knowledge of data cleansing, normalization, metadata management, document parsing, indexing, lineage, version control, and auditability.
  • Familiarity with data modeling for complex scientific, clinical, regulatory, or healthcare knowledge domains.
  • Ability to convert domain-expert needs into practical data structures, metadata models, retrieval-ready content, and maintainable pipelines.
  • Exceptional attention to detail and a methodical approach to identifying data-quality problems.
  • Ability to work with AI engineers, product managers, COA scientists, software engineers, security stakeholders, legal teams, and commercial teams.
  • Strong technical documentation skills.
  • Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management is strongly preferred.
  • Experience with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, medical knowledge repositories, or comparable evidence sources.
  • Experience with vector databases, embeddings, semantic search, Elasticsearch or OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies.
  • Experience using cloud data platforms such as Azure, AWS, or GCP.
  • Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy design, or controlled vocabularies is desirable.
  • Familiarity with ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation.
  • Understanding of GDPR, data privacy, intellectual property restrictions, licensed content, confidential client data, access controls, and secure data handling.
  • Ability to work independently in a remote or hybrid setting while collaborating with global teams.
  • Substantial experience using AI tools in professional work and fluency in English.
  • benefits:[

Benefits

  • A rewarding, progressive career path.
  • Training and ongoing support.
  • Opportunities for professional development and growth.
  • Meaningful influence over the development and delivery of innovative solutions.
  • A multicultural, collegial, and collaborative working environment.

Related Jobs

HarperCollins Publishers UK

Editorial Assistant, Fantasy and Science Fiction

HarperCollins Publishers UK
1,001 – 5,000 Employees
eCommerceMediaRetail

Support HarperCollins’ HarperVoyager and HarperMagpie imprints with editorial administration and title production. The role spans metadata, proofreading, author and agent liaison, and publishing workflows.

Open