In Person

BSides Toronto 2026

10:00 am
PST
October 3, 2026
Toronto Metropolitan University, Ted Rogers School of Management

Thousands of commits hit public GitHub repositories every minute, and a meaningful slice of them are hostile: credential stealers, reverse shells, crypto drainers, and nation-state lures wearing the costume of a coding challenge. The same things that make GitHub great for developers (openness, trust, free hosting, a domain nobody blocks) make it excellent disposable infrastructure for attackers.
This talk is about what happens when you try to watch all of it. I'll walk through a pipeline that scans the public event stream in near real time, the deobfuscation engine that turns walls of XOR'd, packed, and base64'd gibberish back into something a detection rule can match, and the messy reality of keeping false positives low enough that a human analyst still trusts the queue. Then the fun part: who's actually out there, including DPRK-aligned crews running fake job interviews to backdoor developers at crypto firms. Much of the data is first triaged by an autonomous AI analyst turned loose on live adversary infrastructure from a throwaway VM.
You'll leave knowing how to build this visibility yourself, and why GitHub belongs in your threat model next to email and the browser.

GitHub isn't just where developers work. It's where adversaries stage, obfuscate, and deliver malicious code. Every minute, thousands of commits hit public repositories, and buried inside that firehose are credential stealers, reverse shells, crypto drainers, and the occasional nation-state lure dressed up as a coding challenge. The platform's openness, trust, and sheer volume are exactly what make it useful to attackers: free hosting, free CDN, a developer-friendly domain in every allowlist, and a culture where running npm install or cloning a stranger's repo is just Tuesday.

This talk is about what happens when you actually try to watch all of it.

We'll walk through github-threat-scanner, a pipeline that consumes the GitHub public event stream in near real time, pulls down the code behind every push, and runs it through a stack of decoders and detection rules looking for anything that smells wrong. The interesting problems aren't where you'd expect. Ingesting the stream is easy. Storing it is a solved problem. The hard parts are everything in between: peeling back the layers of obfuscation attackers use to hide payloads, deciding what "malicious" even means when half the internet's legitimate code looks suspicious, and keeping false positives low enough that a human analyst can still trust the queue.

We'll dig into the deobfuscation engine (CyberSaucier), a library of CyberChef recipes that chain together XOR bruteforcing, base64 and hex decoding, packed-JavaScript unwrapping, PowerShell de-munging, and the other tricks that turn a wall of gibberish back into something a detection rule can match on. You'll see which recipes earn their keep, which ones we retired because they were pure theatre, and the surprisingly mundane reasons some decoders fail in production that never show up in a blog post.

Then we'll get to the fun part: who's actually out there. Commodity and Nation State actors treat GitHub Pages as disposable infrastructure. And threading through all of it are the targeted operations: DPRK-aligned clusters running fake job interviews and "technical assessments" that ship trojanized projects to developers at crypto firms and long-running personas that maintain plausible commit histories for months before turning hostile.

You'll leave with a concrete picture of how to build this kind of visibility yourself, what the detection surface actually looks like once you're watching it, and why GitHub deserves a seat in your threat model next to email and the browser. If you run a security team, you'll walk out with questions to take back to your developers. If you write detections, you'll have new ideas for where to point them. And if you just like watching adversaries do dumb things at scale, there will be plenty of that too.

The best part of all of this? Most of this data was initially triaged and analyzed by an autonomous AI analyst running in a throwaway VM in dangerous mode, unafraid of touching actual adversary infrastructure.

No prior knowledge of GitHub internals required. Bring opinions about regex.

‍