The GitHub Universe 2026 schedule runs longer than any two-day agenda can hold. This list is built around three questions worth carrying into the conference: how do we know an agent's code actually works, what should an agent remember, and what are we trusting when we install a dependency? That puts memory, evaluations and permissions near the top, with room left for JavaScript tooling and software built for unreliable networks.
Dependencies, provenance and the npm install path
Karen Li and Leo Balter from GitHub trace the dependencies pulled in by npm install back through the systems that publish and protect them, covering the people, permissions and release processes most of us never examine. The session looks at what npm audit cannot catch and where package provenance and OpenID Connect fit in.
Agent memory and the cost of accumulated context
GitHub researchers Cooper Nederhood and Alejandro Carderera built a benchmark from sequences of real pull requests and found that accumulated context can hurt performance. The talk covers how that research shapes their work on agent memory, including features still in development — useful for anyone deciding not just what context to hand an agent, but what to withhold.
Context as shared infrastructure
Instructions, skills and MCP servers that work for one developer do not automatically work for a team. Christopher Harrison from GitHub breaks down what each tool is good for and how to distribute context consistently as a setup grows.
Permissions for hosted MCP servers
"Open a pull request, but do not merge" is a boundary better enforced than written into an instructions file. Nick Taylor from Pomerium demonstrates an identity-aware proxy that adds per-identity authorization in front of a hosted MCP server without modifying the upstream server — and how the rule holds when an agent tries to act on it.
Evaluations and verification
Strong benchmark scores do not guarantee that a model will satisfy the developers using it. Walker Chabbott from GitHub and Julia Kasper from Microsoft show how Copilot evaluates models in production, including what the team measures and what it stopped measuring. The session is a Sandbox Session, built around the question of what makes an evaluation useful.
Verification is the second half of that problem. Jeff An from Momentic covers how agents investigate applications, reproduce unexpected behavior, and separate product bugs from broken tests or infrastructure failures — with attention to the deterministic controls that bound what those agents can do.
Architectures that keep the model on a short leash
In the root-cause analysis architecture Achin Gupta from Intuit and Divya Mahajan from Amazon present, deterministic code handles signal collection, topology traversal, correlation and scoring, while the language model narrates the evidence. The split keeps the agent investigating rather than hallucinating, and the reasoning behind it is worth hearing directly.
Pipeline and toolchain work
Steve Glass and Greg Ose from GitHub map supply chain attack techniques to GitHub Actions controls, covering the ecosystem, the workflow attack surface and runner infrastructure. A workflow holding credentials and release permissions is an attractive target, so the framework is aimed at where defenses belong in a pipeline trusted with a release.
On the tooling side, Alexander Lichter from VoidZero walks through Vite+, an open source CLI for managing the front-end toolchain — bundling, testing, linting, formatting and runtime management included — in the context of a real migration.
Building where the network isn't
Alex Junior Antwi from Braveon AI shares lessons from building CarbonSight for low-connectivity communities in Ghana, with an offline demo. The argument for building for your least-connected users is a familiar one; the implementation details are less so.
A hands-on skill workshop
Shishir Tewari from Procore joins Build once, run on any agent: a practical guide to agent skills, aimed at when a workflow belongs in a reusable skill. The session pairs skills built for recurring personal work with experience building skills for a production data engineering team.



