Skip to content

Improve PoC collection using GitHub archive data #2429

Description

@ziadhany

We are currently collecting PoCs primarily from this repository:

From my understanding, this repository is generated by an automated bot.

There is another project doing something very similar:

The general approach seems to be running a CI job that uses the GitHub API to search for repositories containing CVE IDs and then collecting the results.

However, I think we could use a cleaner and potentially more reliable approach for collecting PoCs.

Instead of repeatedly querying the GitHub API for every CVE ID, we could use the GitHub hourly archive data:

The idea would be to process the archive data locally and search for CVE IDs across newly indexed GitHub content. This could significantly reduce the number of GitHub API requests and give us a more reproducible dataset.

We could then add an extra validation layer to determine whether a discovered repository is actually a valid PoC/Exploit repository. For example, this could involve:

  • Manual review by contributors for higher-confidence results.
  • An LLM-based validation step that reads the repository metadata/content and determines whether it actually contains a PoC or exploit related to the identified CVE.
  • Potentially combining both approaches to assign a confidence level to each result.

As a proof of concept, I’ve tested this approach with:

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions