Data Engineer (ETL)Mid-level ● RemoteFull-timeAn hour ago

Career Path Advertising
Do You Want to Discover Your Ideal Career Path?Craft Your Own Career Path for Free!

About the job

The Middle Data Engineer will build the data foundation for some internal applications, including vendor-file ingestion and normalization, application data structures, data-quality controls, operational reporting data, and AI-assisted vendor-report extraction.

This role will partner closely with the Full Stack Software Engineer, product manager, Stability users, BI Analyst, QA Engineer, and shared IT platform resources. The role is responsible for turning vendor-provided stability information into reviewable, traceable operational data without replacing authoritative GxP systems.

What will you do?

  • Design and implement ingestion and normalization processes for vendor stability reports and related files.
  • Build data structures supporting programs, studies, lots, time points, assay attributes, results, specifications, deviations, vendors, reference standards, and critical reagents.
  • Support automatic generation and tracking of operational time points using study start dates and applicable scheduling rules.
  • Implement data-quality controls for completeness, consistency, valid relationships, duplicate information, and reliable status tracking.
  • Build the data-processing workflow for AI-assisted vendor-report extraction, including proposed values, confidence indicators, human review, and approved value retention.
  • Preserve source references and traceability for extracted or transcribed vendor-derived information.
  • Ensure extracted values are never saved automatically and that AI-assisted functionality remains limited to data entry support.
  • Prepare operational data for portfolio views, overdue-report tracking, review status, OOS/OOT visibility, deviations, trends, and CSV exports.
  • Collaborate with the BI Analyst on operational metrics, dashboard data, trend visualization, and reporting validation.
  • Build pipeline monitoring, error handling, reconciliation, documentation, and support procedures.
  • Partner with the Full Stack Engineer on application data models, APIs, workflow state, and data-quality behavior.
  • Participate in backlog refinement, technical design, testing, release planning, production support, and knowledge transfer.
  • Help establish reusable data-engineering practices for future ITJ TechOps products without expanding our internal application beyond its approved scope.

Qualifications

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, Engineering, or a related field, or equivalent practical experience.
  • 3-4 years of professional data engineering experience delivering production pipelines or data products.
  • Strong Python and SQL skills.
  • Experience with data ingestion, ETL/ELT, transformation, data modeling, orchestration, and data-quality validation.
  • Experience working with cloud data platforms, relational databases, and analytical data stores.
  • Experience building production data processes with logging, monitoring, error handling, retry behavior, and operational support.
  • Ability to work with business and scientific users to understand source files, definitions, workflows, and data-quality expectations.
  • Strong communication, documentation, and cross-functional collaboration skills.

Preferred Qualifications

  • Experience with Databricks, Spark, Delta Lake, Unity Catalog, or comparable technologies.
  • Experience extracting structured data from PDFs or semi-structured vendor files.
  • Experience with document intelligence, OCR, AI-assisted extraction, confidence scoring, or human-in-the-loop data review.
  • Experience with pharmaceutical stability, analytical development, CMC, laboratory, LIMS, or related life-sciences data.
  • Familiarity with specifications, assay results, OOS/OOT indicators, deviations, investigations, and stability trends.
  • Experience designing governed data products for dashboards, analytics, or operational applications.
  • Familiarity with non-GxP/GxP boundaries and source-system ownership.

Skills

Hard Skills

Data Integration

Soft Skills

Analysis and Problem SolvingAgile MindsetCommunication ProficiencyCollaboration

Technical Expertise

Data Engineering