The Biggest Dataset Nobody Modeled: Building an Open Semantic Layer for Telemetry with DuckDB and DataFusion

November 2, 12:20 PM–12:45 PM (PST) • Room 1

Telemetry, the petabytes of metrics, logs, and traces your systems emit, is one of the largest datasets most companies own and one of the last that analytics engineering teams have never modeled. It’s messy, you usually can’t query it with SQL, and it often lives in a silo that requires "SRE" in your title to get full access.

Open standards and open-source projects like DuckDB and DataFusion are changing that: telemetry can now be stored cheaply and queried with regular SQL on open table formats. The hard part is modeling it. Production emits thousands of undocumented event types that vary across vendors. Nobody is going to hand-write a YAML semantic layer for that. However, new OpenTelemetry standards help enormously, which also mean you and your agents can reason about observability regardless of where the data lives.

This talk shows how to infer a vendor-neutral semantuc layer for telemetry designed for AI agents on open table formats. We'll share real numbers, including how agents perform with and without the model, and be honest about where it breaks and when you shouldn't throw a petabyte of logs into an S3 bucket.

OSA CON logo

Join OSA Con 2026

The anti-hype conference for open-source analytics and AI. Connect with engineers, maintainers, and technology leaders building the future of data.

November 2, 2026
San Francisco + Online

Register