01 / Data foundations

Raw data vs views: keep the evidence, shape the answer

Design a useful view of business data without confusing a filtered answer with the underlying evidence. No database experience required.

Reviewed

Three layers with different jobs

A source record is what a business system knows about an item: an issue, an order or a document. Raw data usually means keeping that record close to its original form. This gives you something to revisit when a definition changes or a summary looks wrong. Raw does not mean complete, current or correct; it only describes how little transformation has happened.

An imported table is already a choice. A connector may select fields, convert types and leave out unsupported objects. A GitHub issue table is not a backup of everything in GitHub. Before relying on it, ask which fields are retained, what one row represents and which source objects are missing.

A view is a reusable way to select or transform data for a task. For example, it might expose only open issues and a few useful columns. A database view usually stores the query rather than a separate copy of its results. A materialized view stores results and needs its own refresh process. Neither makes the underlying source fresher.

Work backwards from one question

Suppose an operations lead asks: which open issues should we discuss at the weekly planning meeting? Start with the imported issue records, not a generated summary. Define the candidate set as issues whose state was open at the last successful import. Keep the issue number, title, source link and last-updated timestamp so a reviewer can inspect each candidate.

Then choose an explicit rule, such as open issues with no update for fourteen days. That is a queue for review, not a measure of priority or proof that work has stopped. A conversation could be happening elsewhere, or the issue could intentionally be waiting. Naming the output ‘open issues needing a freshness review’ makes that distinction visible.

Illustrative definition — not a Lake saved-view feature

Question: What should we review at planning?
Grain: One row per imported GitHub issue
Filter: state = open; updated_at before review cutoff
Keep: number, title, URL, updated_at
Freshness: last successful manual sync
Decision: a person reviews the source before acting

Write down the definition before sharing the result

A useful view is a small agreement between the person asking and the person answering. Record its purpose, row grain, filters and time basis. Distinguish source time, such as when an issue changed, from import time, such as when your connector observed it. A new import timestamp does not mean the underlying work is new.

Test the rule with a few records you already understand. Include an item that should qualify, one that should not and one with a missing value. Decide how null timestamps, duplicate records and changed source states should behave. This catches disagreements that a plausible-looking total can hide.

  • Does every row represent the same kind of thing?
  • Are dates interpreted in one stated timezone and interval?
  • Can a reader get back to the original evidence?
  • Are missing values distinguishable from zero or false?
  • Does the output state when and how its source was refreshed?

Avoid the two most expensive shortcuts

First, do not treat a convenient slice as the whole business. An issues table cannot explain customer sentiment, review discussions or deployment health unless those records are actually available. State the coverage alongside the conclusion. ‘No matching rows in this imported dataset’ is more defensible than ‘nothing needs attention.’

Second, do not assume that hiding columns is access control. A view is only a security boundary when the surrounding permissions enforce it. If a client can still read the underlying table, a narrower display does not remove that access. Choose permissions separately from presentation.

How this maps to Lake today

Lake currently imports selected GitHub repository datasets into connector-defined tables. Its read-only MCP tools let an authorised client discover tables, inspect columns and select, filter, sort and page through rows. That is enough to build a task-specific result during an investigation.

Saved views, custom columns and arbitrary SQL joins or aggregations are deferred. The current connector normalizes selected fields; it is not a full raw-payload archive. Use the distinction in this guide to plan useful questions now, and treat reusable database views as a concept rather than an available Lake setup step.