Zach joins Kostas and Nitay to trace how security data turned into a full-blown data engineering problem. He walks through his path from hospital networks with a sprawling attack surface, into tightly regulated banking, and then into big tech, where petabyte-scale volumes broke the Splunk-style stack he was hired to optimize and pushed him toward columnar storage and table formats like Iceberg and Delta.
From there the conversation digs into OCSF, the Open Cybersecurity Schema Framework: why a producer-side schema would shrink the mountain of cleanup work defenders do today, how an unusually open community formed around it in a famously secretive industry, and why security schemas have to standardize semantics and values, not just structure. Zach, Kostas and Nitay also get into the overlap between security and observability stacks, the unknown-unknowns problem that forces teams to store data they may never query, the case for a metrics-plus-security semantic layer, and what changes when agents both need clean context and become the thing you have to monitor.
Chapters00:00 Introduction and Zach's background in cybersecurity
01:09 Early days in healthcare cybersecurity and data detection
02:19 The rise of Splunk and schema on read in security
03:07 Moving from healthcare to banking and regulated environments
04:21 Transition to big tech and handling massive data volumes
08:35 How Splunk's distributed compute works and scaling challenges
10:45 The shift towards modern data stacks and table formats
11:32 Introduction to the Open Cyber Security Schema Framework (OCSF)
14:00 Community and collaboration in cybersecurity data standards
15:36 The future of data schemas and standardization in security
20:59 Behavioral analytics and threat modeling in cybersecurity
23:28 The role of schemas in reducing security data complexity
28:01 Convergence of security data and observability
34:46 The impact of LLMs and AI on cybersecurity practices
50:46 The future of OCSF and standardization in cybersecurity