Cedric Renggli

cedric_renggli.jpg

I am a Senior Researcher and Lecturer at ETH’s Systems Group working with Ana Klimovic.

My research lies at the intersection of data management and AI systems. I research and build systems that make emerging AI workloads efficient, adaptive, and measurable. My work addresses three connected questions: How should AI systems manage the data and knowledge they consume and produce? How can stateful AI workloads be executed efficiently? And how can their behavior be measured reliably from noisy observations?

Previously, I was a Senior Researcher at Apple, a PostDoc at UZH (DaST with Dan Olteanu) and defended my thesis at ETH’s Systems Group with Ce Zhang.

If you are interested in collaborating, please reach out directly. I also welcome motivated students at ETH for projects and theses. To apply, send me an email with your CV and transcript of records attached.

Research Focus Areas

I develop data systems for efficiently organizing, maintaining, and querying the data used by AI applications. My previous work includes scalable search over pretrained models in SHiFT, efficient in-database learning without full data shuffling in CorgiPile, and incremental maintenance of vector-search indexes in Ada-IVF. Building on these foundations, I am interested in systems that automatically transform evolving data, model outputs, and interaction traces into maintainable and queryable knowledge, with statistical guarantees on answer quality and uncertainty.

Emerging AI workloads are increasingly stateful and interactive, comprising sequences of model calls, retrieval operations, tool executions, and user interventions whose resource demands evolve over time. I design serving architectures, state-management mechanisms, and adaptive scheduling policies for these workloads. This agenda extends my earlier work on communication-efficient distributed learning in SparCML and robust execution across unreliable infrastructure in Distributed Learning over Unreliable Networks. My current interests include managing intermediate execution state (e.g., KV caches), predicting future resource requirements from interaction traces, and coordinating model inference and secure tool execution across heterogeneous hardware.

The observed performance of an AI system reflects not only model capability, but also data quality, benchmark construction, infrastructure, execution environments, and stochastic variation. I develop methods for identifying and quantifying these effects. My work has exposed systematic ambiguities and metric biases in Text-to-SQL evaluation, introduced automatic feasibility assessment through data-quality analysis in Snoopy, and established statistically rigorous testing methods for ML pipelines in ease.ml/ci. By disentangling model behavior from data, system, and measurement artifacts, this research enables more reliable comparisons and better-informed decisions about complete AI systems.

Selected Publications

  1. Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
    Cedric Renggli, Ihab F Ilyas, and Theodoros Rekatsinas
    arXiv preprint arXiv:2501.18197, 2025
  2. SHiFT: an efficient, flexible search engine for transfer learning
    Cedric Renggli, Xiaozhe Yao, Luka Kolar, and 3 more authors
    Proceedings of the VLDB Endowment, 2022
  3. Automatic feasibility study via data quality analysis for ml: A case-study on label noise
    Cedric Renggli, Luka Rimanic, Luka Kolar, and 2 more authors
    2023 IEEE 39th International Conference on Data Engineering (ICDE), 2023
  4. A Data Quality-Driven View of MLOps
    Cedric Renggli, Luka Rimanic, Nezihe Merve Gürel, and 3 more authors
    IEEE Data Engineering Bulletin, 2021
  5. Continuous Integration of Machine Learning Models with ease.ML/CI: Towards a Rigorous Yet Practical Treatment
    Cedric Renggli, Bojan Karlas, Bolin Ding, and 4 more authors
    In SysML Conference, 2019