Start with a successful import
In Lake's Connections area, add a GitHub Repo connection with the repository owner/name, a unique table prefix and a fine-grained personal access token restricted to that repository. Choose the datasets needed for your questions: commits, issues, pull requests or GitHub Actions workflow runs. The connector supports github.com; commits follow the default branch.
Use read-only repository permissions for the selected datasets: Contents for commits, Issues for issues, Pull requests for pull requests and Actions for workflow runs. Metadata read access is also needed, and an organisation may require token approval. Run a manual sync, check its result and inspect the imported tables in Data viewer before asking an agent to interpret them.
Connecting creates tables but does not by itself populate them. All four tables exist even if you select fewer datasets, so an empty table may simply be unselected. Sync is manual, not scheduled. A failed page or rate limit rolls back that run's imported rows; check the last successful sync rather than assuming the latest attempt refreshed the data.
1. Which open issues should we revisit?
Ask for imported open issues ordered by their source update time, oldest first. Include the issue number, title and source URL. Choose a review cutoff that makes sense for the team, and describe the result as a review queue rather than a list of neglected work. An old timestamp is a useful prompt for a conversation, not proof of a problem.
Lake excludes pull requests returned by GitHub's issues endpoint from the issues table. That makes the row meaning clearer, but it does not supply every bit of issue context. Do not assume comments, project-board state or linked customer requests are present; inspect the schema and follow the source link for details.
2. Which open pull requests need a closer look?
Request open pull requests with their draft status, update time and source link. Separate drafts from non-drafts, then ask a person to review the oldest candidates. The imported draft field uses integers: zero means false and one means true. Confirm field names with describe_table instead of guessing them from the display labels.
An open, non-draft pull request is not necessarily ready to merge. Review approvals, branch protections, conflicts and linked CI checks require evidence beyond that status. Do not turn this queue into a merge recommendation unless you have checked those conditions in GitHub.
3. Which failed workflow runs deserve investigation?
Inspect recent GitHub Actions workflow runs with a failure conclusion and include their source URLs. Separate a run's execution status from its conclusion: an in-progress run has not yet produced a final outcome. Choose an explicit date window and inspect the corresponding columns before filtering.
These rows represent workflow runs, not individual jobs, logs, artifacts or third-party checks. Several failures could be retries of the same underlying problem. Use the queue to open GitHub and investigate; the imported run metadata alone cannot explain a stack trace or establish a root cause.
Use a schema-first query workflow
Connect a compatible MCP client using a workspace-issued bearer credential. First call list_tables, then describe_table on the relevant table. Use query_table to select only useful columns and apply supported filters and ordering. Table names include your chosen prefix, so never copy another workspace's table name blindly.
Each query returns at most one hundred rows. rowCount is the count for that page, not a total count. Continue with nextOffset while additional pages are available; if hasMore is true but nextOffset is null, narrow the filters because the offset ceiling has been reached. Keep the same sorting and filters between pages, and remember that changing data can shift offset-based results.
Ask your agent Discover the schema for my imported GitHub issues. Build a review queue of open issues before my cutoff. Keep the source URL and update timestamp. Use stable sorting and check for additional pages. State whether the result is complete or a sample. Report missing context before suggesting follow-up.
Keep the conclusion narrower than the evidence
Previously imported rows remain after upstream deletion, force-pushes or dataset deselection. That means an import is not a live mirror, and absence or presence in the table needs interpretation. Verify important records in GitHub before changing a plan.
Lake's current MCP queries are structured single-table reads. They do not provide arbitrary SQL, joins or server-side aggregations. If you calculate a total from retrieved pages, label it as a calculation over that retrieved set and document completeness. These three questions are useful precisely because they produce inspectable queues without pretending to measure the health of an entire engineering organisation.
