We are currently collecting PoCs primarily from this repository:
From my understanding, this repository is generated by an automated bot.
There is another project doing something very similar:
The general approach seems to be running a CI job that uses the GitHub API to search for repositories containing CVE IDs and then collecting the results.
However, I think we could use a cleaner and potentially more reliable approach for collecting PoCs.
Instead of repeatedly querying the GitHub API for every CVE ID, we could use the GitHub hourly archive data:
The idea would be to process the archive data locally and search for CVE IDs across newly indexed GitHub content. This could significantly reduce the number of GitHub API requests and give us a more reproducible dataset.
We could then add an extra validation layer to determine whether a discovered repository is actually a valid PoC/Exploit repository. For example, this could involve:
- Manual review by contributors for higher-confidence results.
- An LLM-based validation step that reads the repository metadata/content and determines whether it actually contains a PoC or exploit related to the identified CVE.
- Potentially combining both approaches to assign a confidence level to each result.
As a proof of concept, I’ve tested this approach with:
We are currently collecting PoCs primarily from this repository:
From my understanding, this repository is generated by an automated bot.
There is another project doing something very similar:
The general approach seems to be running a CI job that uses the GitHub API to search for repositories containing CVE IDs and then collecting the results.
However, I think we could use a cleaner and potentially more reliable approach for collecting PoCs.
Instead of repeatedly querying the GitHub API for every CVE ID, we could use the GitHub hourly archive data:
The idea would be to process the archive data locally and search for CVE IDs across newly indexed GitHub content. This could significantly reduce the number of GitHub API requests and give us a more reproducible dataset.
We could then add an extra validation layer to determine whether a discovered repository is actually a valid PoC/Exploit repository. For example, this could involve:
As a proof of concept, I’ve tested this approach with: