Citylitics Inc.
Citylitics Inc.
Citylitics Inc. combines artificial intelligence, consulting expertise, and large-scale data analysis to help leaders in the infrastructure industry understand North American local utility and public infrastructure markets. Its Data Engine draws on information from more than 31,000 city and utility sources, turning complex market signals into practical insight about where infrastructure investment is emerging and how projects may develop. The company supports more informed business decisions across a sector facing substantial long-term investment needs.

Data Engineer – AI Pipelines and Data Products

Build production data pipelines and agentic AI workflows for Citylitics’ infrastructure intelligence platform. Own data products from orchestration through customer-facing applications.

Description

  • Build high-performance, reliable, maintainable data pipelines with Airflow and BigQuery
  • Lead pipeline architecture across design, testing, deployment, monitoring, and maintenance
  • Make independent technical decisions and operate the systems you build
  • Work with analysts and stakeholders to define requirements and create scalable data models and applications
  • Manage data products end to end, from pipelines and models to backend APIs, dashboards, and applications
  • Design and run AI and agentic workflows that convert raw documents into structured data
  • Create prompts and context strategies using tool and function calling, embeddings, retrieval, and structured outputs
  • Address production requirements such as retries, idempotency, cost control, and latency
  • Maintain evaluation datasets, LLM-as-judge assessments, rule-based quality checks, and benchmarks
  • Apply evaluation findings to model, prompt, and architecture decisions
  • Improve data infrastructure and engineering processes continuously
  • Assess and adopt relevant technologies and engineering practices
  • Help define how AI coding agents support development, debugging, code review, and documentation
  • Handle additional responsibilities as assigned

Requirements

  • Proficiency in Python and SQL
  • At least three years of experience as a Data Engineer or Software Engineer in a cloud environment
  • Experience building production data pipelines with Airflow or Cloud Composer and BigQuery
  • Knowledge of data modeling, ETL/ELT, and reliable idempotent pipeline design
  • Experience delivering LLM-powered production features or pipelines with APIs such as Gemini, Claude, or OpenAI
  • Experience with prompting, structured outputs, and tool or function calling
  • Working knowledge of agent patterns, including tool use, multi-step workflows, retrieval/RAG, and MCP
  • Understanding of production tradeoffs involving cost, latency, and failure modes
  • Daily fluent use of AI coding agents such as Claude Code or Cursor, including reviewing and verifying generated output
  • Experience with Git, CI/CD, and cloud platforms; GCP experience is preferred
  • Experience developing end-to-end data applications or dashboards with backend APIs and frontends such as React and TypeScript is a plus
  • Strong problem-solving, independent working, communication, and collaboration skills

Benefits

  • Influence sustainable public infrastructure through work with real-world impact
  • Represent a differentiated, data-driven infrastructure platform
  • Work across infrastructure, scale-up operations, and data science
  • Join a fast-paced environment with limited corporate bureaucracy
  • Use generative AI tools and access the company’s full data universe
  • Receive in-role coaching, skills-based development, and clear internal promotion pathways
  • Work in a collaborative team that shares celebrations and support
  • Join a safe, diverse, and inclusive workplace
  • Work for an equal opportunity employer

Related Jobs

The Clorox Company

Shopper Insights & Analytics Manager - Amazon, Hybrid

The Clorox Company

Lead shopper and category analytics that support The Clorox Company’s Amazon growth. Turn behavior, research, and retail data into decisions on assortment, pricing, content, conversion, and repeat purchasing.

Open