EVEREST

Everest Lab logo

Data Systems Group · MIT CSAIL

Our mission is to establish the engineering principles for AI-driven data systems. We move beyond ad-hoc experimentation to build a foundation of new abstractions, reusable software components, and rigorous development tools.

Research Directions

Code Generation

Developing AI systems that automatically generate code from specifications, design documents, and natural language descriptions, advancing program synthesis and software development automation.

Data Analytics

Building AI-powered systems for data analysis, transformation, and insight extraction, applying optimization techniques to make data processing more efficient and accessible.

Large Scale Retrieval

Developing efficient systems for retrieving relevant information from massive datasets, optimizing both accuracy and performance for real-world applications.

Agent Operations

Investigating operational frameworks for deploying and managing AI agents in production environments, addressing challenges in monitoring, debugging, and maintaining autonomous systems.

Agent Planning

Developing methods for AI agents to autonomously plan and execute complex multi-step tasks, combining reasoning, decision-making, and adaptive execution strategies.

Projects

G5: Design Doc Driven Development

G5 achieves a 54.8% average pass rate on the commit-0 benchmark, outperforming the next-best system by 14 percentage points. The system is an LLM-based compiler that generates code from design documents, treating design specifications as input and automatically producing implementation code.

Palimpzest

Palimpzest is a declarative system for AI-powered analytics over unstructured data. Its cost-based optimizer chooses models, prompts, and execution strategies to balance runtime, cost, and output quality. With parallelism, it found plans up to 90.3x faster and 9.1x cheaper than a single-threaded GPT-4 baseline, with an F1 score within 83.5% of that baseline.

palimpzest.org ↗

Carnot

Carnot is an interactive execution engine for AI-driven analytics, presented as a VLDB 2026 demo. It compiles natural-language requests into execution graphs shown in a notebook, where analysts can inspect intermediate data, edit operators, and set cost or latency limits for its query optimizer.

github.com ↗

KramaBench

KramaBench is a benchmark for AI systems that build data-to-insight pipelines over data lakes, presented at ICLR 2026. Its 104 manually curated challenges span 1,700 files, 24 data sources, and 6 domains, and test whether a system can orchestrate extraction, cleaning, integration, analysis, and modeling end to end.

kramabench.org ↗

Log-Augmented Generation

Log-augmented generation (LAG) lets language models reuse reasoning from earlier tasks. It stores logs of past tasks as key-value caches and retrieves the relevant ones for each new task, outperforming agentic systems that do not use logs on knowledge- and reasoning-intensive datasets.

peterbaile.github.io ↗

Archi

Archi is an open-source framework for scientific collaborations: it organizes heterogeneous data sources and deploys agents that retrieve and reason over them. It supports MIT Physics' SubMIT computing cluster and MIT courses, and since February 2026 it has served as a support agent for the CMS experiment's computing operations team at CERN.

github.com ↗

BRAD

BRAD is a cloud data virtualization system. Users query one SQL interface backed by multiple cloud database services, and BRAD picks the best engine for each query, provisions resources to minimize cost, and adapts as workloads shift. In the authors' evaluation it met performance targets while saving 1.6–13x in cost compared with serverless auto-scaling or HTAP systems.

dsg.csail.mit.edu ↗

Team

Faculty

Samuel Madden

Tim Kraska

Michael Cafarella

Omar Khattab

Affiliate Faculty

Christoph Paus

Postdocs

Jason Mohoney

Postdoctoral Researcher

Gerardo Vitagliano

Postdoctoral Researcher

Geoffrey Yu

Postdoctoral Researcher

Xinjing Zhou

Postdoctoral Researcher

Students

Peter Baile Chen

PhD Student

Zhuohan (Joshua) Gu

PhD Student

Darryl Ho

PhD Student

Eugenie Lai

PhD Student

Amadou Latyr Ngom

PhD Student

Matthew Russo

PhD Student

Anna Zeng

PhD Student

Sylvia Zhang

PhD Student

Alumni

Ferdinand Kossmann

Tianyu Li

Markos Markakis

Ziniu Wu

Publications

Recent papers by Everest members, drawn from DBLP. Current and former members are shown in bold.

  1. Ken: An Execution Engine for Unstructured Database Systems

    Ferdinand Kossmann, Ziniu Wu, Alex Turk, Nesime Tatbul, Lei Cao, Samuel Madden

    Proc. VLDB Endow. 2026

  2. Abacus: A Cost-Based Optimizer for Semantic Operator Systems

    Matthew Russo, Chunwei Liu, Sivaprasad Sudhir, Gerardo Vitagliano, Michael J. Cafarella, Tim Kraska, Samuel Madden

    Proc. VLDB Endow. 2026

  3. Consistency and Correctness in Data-Oriented Workflow Systems

    Michael Stonebraker, Xinjing Zhou, Peter Kraft, Qian Li

    CIDR 2026

  4. Making Array-Based Translation Practical for Modern, High-Performance Buffer Management

    Xinjing Zhou, Jinming Hu, Andrew Pavlo, Michael Stonebraker

    arXiv 2026

  5. Deep Research is the New Analytics System: Towards Building the Runtime for AI-Driven Analytics

    Matthew Russo, Tim Kraska

    CIDR 2026

  6. Improving DBMS Scheduling Decisions with Accurate Performance Prediction on Concurrent Queries

    Ziniu Wu, Markos Markakis, Chunwei Liu, Peter Baile Chen, Balakrishnan Narayanaswamy, Tim Kraska, Samuel Madden

    Proc. VLDB Endow. 2025

  7. Palimpzest: Optimizing AI-Powered Analytics with Declarative Query Processing

    Chunwei Liu, Matthew Russo, Michael J. Cafarella, Lei Cao, Peter Baile Chen, Zui Chen, Michael J. Franklin, Tim Kraska, Samuel Madden, Rana Shahout, Gerardo Vitagliano

    CIDR 2025

  8. DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification

    Darryl Ho, Samuel Madden

    CVPR 2025

Show 17 moreShow fewer
  1. Virtualizing Cloud Data Infrastructures with BRAD

    Geoffrey X. Yu, Ziniu Wu, Ferdi Kossmann, Tianyu Li, Markos Markakis, Amadou Ngom, Sophie Zhang, Tim Kraska, Samuel Madden

    SIGMOD Conference Companion 2025

  2. Log-Augmented Generation: Scaling Test-Time Reasoning with Reusable Computation

    Peter Baile Chen, Yi Zhang, Dan Roth, Samuel Madden, Jacob Andreas, Michael J. Cafarella

    arXiv 2025

  3. KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes

    Eugenie Lai, Gerardo Vitagliano, Ziyu Zhang, Sivaprasad Sudhir, Om Chabra, Anna Zeng, Anton A. Zabreyko, Chenning Li, Ferdi Kossmann, Jialin Ding, Jun Chen, Markos Markakis, Matthew Russo, Weiyang Wang, Ziniu Wu, Michael J. Cafarella, Lei Cao, Samuel Madden, Tim Kraska

    arXiv 2025

  4. CONCUR: A Framework for Continual Constrained and Unconstrained Routing

    Peter Baile Chen, Weiyue Li, Dan Roth, Michael J. Cafarella, Samuel Madden, Jacob Andreas

    arXiv 2025

  5. Causal DAG Summarization

    Anna Zeng, Michael J. Cafarella, Batya Kenig, Markos Markakis, Brit Youngmann, Babak Salimi

    Proc. VLDB Endow. 2025

  6. Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method

    Peter Baile Chen, Yi Zhang, Mike Cafarella, Dan Roth

    ACL 2025

  7. Toward Standardized Data Preparation: A Bottom-Up Approach

    Eugenie Y. Lai, Yuze Lou, Brit Youngmann, Michael J. Cafarella

    EDBT 2025

  8. CausaLens: A System for Summarizing Causal DAGs

    Noam Chen, Anna Zeng, Michael J. Cafarella, Batya Kenig, Markos Markakis, Oren Mishali, Brit Youngmann, Babak Salimi

    SIGMOD Conference Companion 2025

  9. SeerCuts: Explainable Attribute Discretization

    Eugenie Y. Lai, Inbal Croitoru, Noam Bitton, Ariel Shalem, Brit Youngmann, Sainyam Galhotra, El Kindi Rezig, Michael J. Cafarella

    SIGMOD Conference Companion 2025

  10. PalimpChat: Declarative and Interactive AI analytics

    Chunwei Liu, Gerardo Vitagliano, Brandon Rose, Matthew Printz, David Andrew Samson, Michael J. Cafarella

    SIGMOD Conference Companion 2025

  11. EnrichIndex: Using LLMs to Enrich Retrieval Indices Offline

    Peter Baile Chen, Tomer Wolfson, Michael J. Cafarella, Dan Roth

    arXiv 2025

  12. Practical DB-OS Co-Design with Privileged Kernel Bypass

    Xinjing Zhou, Viktor Leis, Jinming Hu, Xiangyao Yu, Michael Stonebraker

    Proc. ACM Manag. Data 2025

  13. Tux: Efficient Drop-in Networking for Database Systems

    Xinjing Zhou, Viktor Leis, Xiangyao Yu, Michael Stonebraker

    Proc. VLDB Endow. 2025

  14. Tiered-Indexing: Optimizing Access Methods for Skew

    Xinjing Zhou, Xiangpeng Hao, Xiangyao Yu, Michael Stonebraker

    VLDB J. 2025

  15. OLTP Through the Looking Glass 16 Years Later: Communication is theNew Bottleneck

    Xinjing Zhou, Viktor Leis, Xiangyao Yu, Michael Stonebraker

    CIDR 2025

  16. Parachute: Single-Pass Bi-Directional Information Passing

    Mihail Stoian, Andreas Zimmerer, Skander Krid, Amadou Ngom, Jialin Ding, Tim Kraska, Andreas Kipf

    Proc. VLDB Endow. 2025

  17. Recursive Language Models

    Alex L. Zhang, Tim Kraska, Omar Khattab

    arXiv 2025

For all publications from the Data Systems Group, see the DSG publications page.

Sponsors

Our research is made possible through the generous support of industry partners and funding agencies who share our vision for advancing AI-driven data systems.

ARPA-H Advanced Research Projects Agency for Health
United States Air Force
Google
Amazon
Intel
Accenture
Two Sigma
InterSystems
TWG Global

For sponsorship opportunities, please contact us at everest-info@csail.mit.edu

Contact

Location

MIT Computer Science & Artificial Intelligence Laboratory

32 Vassar Street

Cambridge, MA 02139

Prospective Students

We welcome inquiries from prospective PhD students interested in AI-powered data systems. Please reach out to faculty members directly regarding research opportunities.