# From warehouse to data contracts A new approach to our data platform ~10 minutes · then questions Note: [0:15] Set the frame. One idea for the whole talk: the way we get data in doesn't scale with a central team, so we're changing who does the work and giving them a pattern to follow. --- ## A very short history 1. **Data warehouse** - structured, central, trusted. But one team models everything, and every new source waits in their queue. 2. **Data lake** - store anything, cheaply. But no schema, no transactions: the "data swamp". 3. **Lakehouse** - open table formats (Iceberg, Delta, Hudi) bring warehouse guarantees to cheap lake storage. 4. **Data mesh and data contracts** - owners publish their own data under an agreed contract. Note: [1:00] Each step fixed one problem and left another. The last one is about *who does the work*, and that's where this talk lands. Dates if asked: Iceberg started at Netflix in 2017 (Apache top-level 2020). Data mesh article 2019. "Lakehouse" popularised 2020-21 (CIDR 2021 paper). Data contracts take shape from 2021; Open Data Contract Standard 2023. Mesh is about who owns data. Lakehouse is about where it lives. Contracts are how owners agree on it. They're complementary, not successive. --- ## Where we are today  <!-- .element: style="max-height:420px; background:none; border:none; box-shadow:none" --> - Every new data set means learning an unfamiliar structure - Fixes happen **downstream**; the team that owns the data never hears about them - Fine for tens of feeds. **Not for hundreds or thousands.** Note: [1:00] Be honest about what works. The point isn't that it's broken; it's that the bottleneck is people, not technology. Don't claim the Hive tables are slow unless you have a number - we don't. The quality problem (from a fellow architect): when a different team owns a transformed copy of the data, there's usually no feedback to the team that produced the original. Ad hoc fixes and inconsistencies build up in the copy and get harder to maintain, instead of being fixed at the source. -- ### You are here Legacy warehouse → **[ target platform ]** Note: Optional marker slide. One line: we're at the left, the next slide is the right. --- ## Why ownership matters **Without it** A copy owned by another team gets patched downstream. The source stays wrong. **With it** The owner sees how their data is used, and fixes it **once, at the source**. Note: [0:45] Contrast the two worlds in one breath each. Without ownership: a different team owns a transformed version, there's no feedback to the producer, ad hoc fixes and inconsistencies build up and get harder to maintain. With ownership and contracts: teams have context on how their data is consumed, so problems come back to the people who can fix them. Timing: this adds about 45 seconds. If you're over, cut the optional "You are here" and Iceberg-structure slides. --- ## Where we want to be - Domain teams **own and publish** their data - Owners see how their data is used, so problems are fixed **at the source** - A **data contract** describes each feed: a reporting schema the team **stands behind**, not their internal model - A **shared pattern** turns a contract into an Iceberg table - Query with **Athena**; export Parquet for local tools - An **MCP layer** for exploring data in plain English - Redshift becomes **one way to consume** the data, not the centre of it - Existing reports keep running while this is built - **Still to build:** a reporting platform on top of the new data Note: [1:15] The key sentence: "You describe your data; the platform does the plumbing." The ownership point: teams then have context on how their data is consumed, which closes the feedback loop. No Redshift end date, and don't say it's going away: it stays for the medium term, may become a small part of the stack, and some legacy reports still depend on it. Airbyte and DMS ingestion is what's being replaced. Be clear that the reporting layer on top isn't built yet - the feeds come first. --- ## What is Apache Iceberg? An **open table format** for big data on object storage. - A table is a set of Parquet files **plus metadata**, not a folder - The metadata tracks schema, files and **snapshots** - A catalog (we use **Glue**) points to the current state - Any engine that speaks Iceberg can read it Note: [1:30] Analogy: a library catalogue. The books (Parquet files) sit on shelves in S3; the catalogue (metadata) says which books make up the table right now. Change the catalogue entry and everyone sees the new version at once. Created at Netflix to fix the limits of Hive tables. -- ### Inside an Iceberg table  <!-- .element: style="max-height:420px; background:none; border:none; box-shadow:none" --> Note: Optional deeper slide - skip it if you're short of time. Each write creates a new snapshot; readers always see a complete one. --- ## What that gives us 1. **Safe schema changes** - columns tracked by ID; add, rename, drop without breaking old data 2. **No engine lock-in** - Athena, Spark and others over the same tables 3. **Consistent reads** - readers see a whole snapshot, never half-written files Note: [1:15] Tie each one back to the ownership story: (1) teams can change their own data safely, (2) we aren't tied to one engine or vendor, (3) many teams can write while people query. Only mention time travel if you have tested it yourself. We've tested querying Iceberg and it has been fast and worked fine; don't claim a benchmark. --- ## How a feed gets in (the target)  <!-- .element: style="max-height:420px; background:none; border:none; box-shadow:none" --> - The contract holds the schema, description and **cadence** - The same description is published to Glue Note: [1:15] Walk left to right. The skill checks the prerequisites before the PR: a contract exists, it's documented, the documentation reaches Glue. Teams follow the README in the Terraform repo; they should be largely autonomous. Better documentation in Glue also means better answers from the LLM layer later. -- ### An example contract ```text Table orders { id bigint [pk, note: 'Order identifier'] email varchar [note: 'Customer email. pii'] placed_at timestamp [note: 'When the order was placed'] Note: 'Orders placed online. Updated daily.' } ``` Note: This is illustrative only. The `pii` tag in the column note is a *proposed* convention, not an agreed one. Say so if asked. --- ## Access and sensitive data - If data **shouldn't be exposed, it shouldn't be in the contract** - Contracts will mark sensitive columns, so they can be restricted - Today the MCP layer is reachable on the **VPN only** - Next: **OIDC and roles** (proof of concept already works) - Direction: roles → Lake Formation tags → visible columns **Not built yet.** Enforcement is the next step. Note: [0:45] Be upfront. Today we've only noted where PII appears in one feed; nothing masks it. The MCP layer is in production and can read the two running feeds, with no OIDC yet. Lake Formation supports column-level access on Iceberg tables in Athena via LF-tags. Per-user restriction only works if the MCP layer queries as the user (or a role mapped to their group), not one shared role. LF-tags and data filters can't be combined. This is the question most likely to be asked first. --- ## Where we are right now | Working | In progress | Planned | |---|---|---| | Contracts repo, some contracts complete and documented | Terraform pattern (clean-up after proof of concept) | Reporting platform | | **Two feeds** running into Iceberg | Onboarding skill | Sensitive-data handling | | Querying through Athena | Fixes from testing | | | MCP layer in production (VPN only) | OIDC for the MCP layer | | Note: [0:30] Two feeds were chosen because they behave differently: one continuous, one a daily refresh. We started by streaming one and settled on two kinds of batch: a full batch, and a batch over a time window. Feeds are cheap to rebuild: if one is wrong, drop the data and rerun it. Testing by a colleague found gaps; fixes are going in. Don't name the feeds or the colleague on the slide. --- ## What I need from you **This week** - **Try the MCP layer** and tell us what's wrong with it - Give us **feedback** on this approach - Tell us about **feeds you'd like to bring in** **We're doing:** authentication and roles, so access is controlled by who you are. -- ### To onboard a feed 1. A **data contract** (DBML) in the contracts repo 2. A **description** on every table and column, including **cadence** 3. A named **owner** 4. **Sensitive columns marked**, or left out Then: follow the README, run the skill, open a PR. *(Coming - the pattern and skill aren't ready yet.)* Note: [1:15] Keep this short; the ask is small on purpose. Right now the pattern and skill aren't finished, so don't send people to self-onboard; ask for conversations about their feeds, which also tells us what the pattern needs to cover. CI will run the same checks as the skill, so a green PR needs little review. We review closely at first and relax as the process matures. Ownership can be split more finely across teams. We may not own the platform long term, so the pattern has to stand without us. --- ## Questions - Where's the contracts repo and the Terraform README? - What happens to my existing dashboards? - What about sensitive data? Note: Q: "If a warehouse ingests all these reporting schemas, how is that different from a data warehouse?" A: (1) Teams *publish* their data; nobody extracts it from them. (2) The team chooses the schema it stands behind as the contract, rather than a private internal model leaking out as one. (3) The engine is a consumer's choice: Athena, a local SQL engine pulling down only the data sets you want, or a warehouse if you prefer one. Teams also choose what they aggregate, so transforms sit with the owners. Other likely questions and honest answers: (1) Redshift timing - no date; it stays for now, may shrink long term, existing reports are unaffected. (2) PII - not built, tagging is the foundation, contract content is the first control. (3) Who reviews - us first, automation later. (4) Is it fast - tested, works fine, no benchmark claimed. (5) Why not just keep the warehouse - doesn't scale to thousands of feeds through one team.