Skip to content

Repository files navigation

LinkedIn Research Lab

CI Python License: MIT

An isolated, read-only research service and benchmark harness for public LinkedIn company pages.

It normalizes source results, records provenance, preserves raw content, and exposes a small HTTP contract that can be consumed by another application. The project is intentionally stateless: the calling application owns durable storage and downstream analysis.

Important: This is a research and integration component, not a LinkedIn automation or outreach tool. It does not send messages, create connections, like, comment, share, extract cookies, bypass CAPTCHAs, or defeat access controls.

What this project does

  • Fetches public company-page content through replaceable source adapters.
  • Supports Crawl4AI as the primary service engine and self-hosted Jina Reader OSS as an alternative.
  • Keeps Agent-Reach and Camoufox adapters available for controlled experiments.
  • Returns raw content, a content hash, title, timestamp, source, engine, and extensible metadata.
  • Provides a bounded batch pipeline with deduplication, error capture, and run manifests.
  • Offers a FastAPI service for integration with systems such as Netsy Leads.

What it deliberately does not do

  • No credential, cookie, session, or proxy management.
  • No login automation or CAPTCHA bypass.
  • No person scraping or contact verification.
  • No outreach, engagement, connection requests, or profile actions.
  • No claim that public LinkedIn access is reliable at arbitrary volume.

Quick start

Requirements

  • Python 3.11+
  • uv
  • Docker, if you want to run the packaged service and Jina sidecar

Run the tests and CLI

uv sync --all-extras
uv run pytest
uv run ruff check src tests scripts
uv run ruff format --check src tests scripts
uv run linkedin-lab data/input/targets.csv --output data/output/results.jsonl

Run the HTTP service

uv run uvicorn linkedin_lab.service:app --host 127.0.0.1 --port 8082

Health check:

Invoke-WebRequest http://127.0.0.1:8082/health

Research request:

$body = @{ company_name = "OpenAI"; linkedin_url = "https://www.linkedin.com/company/openai/"; engine = "crawl4ai" } | ConvertTo-Json
Invoke-RestMethod -Method Post -Uri http://127.0.0.1:8082/v1/company-research -ContentType "application/json" -Body $body

Run with Docker

docker compose up --build

The service is exposed on http://localhost:8082; the Jina Reader sidecar is internal to the Compose network.

HTTP contract

GET /health

{"status":"ok"}

POST /v1/company-research

Request:

{
  "company_name": "OpenAI",
  "linkedin_url": "https://www.linkedin.com/company/openai/",
  "engine": "crawl4ai",
  "extra": {"request_id": "demo-001"}
}

The response includes status, content, content_sha256, title, source, captured_at, and extra. On provider or access failure, the service returns a structured error result rather than pretending that the page was successfully read.

Architecture

CSV / JSONL targets
        |
        v
bounded batch pipeline ---> source adapter
        |                       |- fixture
        |                       |- Crawl4AI
        |                       |- Jina Reader OSS
        |                       |- Agent-Reach (controlled)
        |                       `- Camoufox (optional)
        v
normalized result + provenance + error
        |
        +--> JSONL / CSV output and run manifest
        `--> HTTP service response

The service does not own application persistence. In the Netsy integration, the worker stores raw snapshots and structured extractions in PostgreSQL, including an additional_data JSON field for future provider-specific information.

Data and privacy

Only use targets and fields that your organization is authorized to process. Before a live run, define the target scope, retention period, deletion owner, allowed fields, and stop conditions. Do not commit credentials, cookies, private exports, or customer/lead lists.

The repository includes synthetic examples and public company URLs for smoke tests. Private or customer-specific input files belong outside Git; .gitignore excludes local environment files, output files, and private Netsy exports.

Read the controlled live-test protocol before enabling an experimental adapter.

Development

Source code lives under src/linkedin_lab; tests are under tests. The project uses:

  • uv for dependency and lockfile management;
  • pytest for tests;
  • ruff for linting and formatting;
  • GitHub Actions for repeatable CI.

Run the full local gate:

uv sync --locked --all-extras
uv run pytest -q
uv run ruff check src tests scripts
uv run ruff format --check src tests scripts

Contributing

Contributions are welcome, especially new source adapters, parser tests, provenance improvements, and reproducible fixtures.

  1. Read CONTRIBUTING.md.
  2. Keep adapters read-only and bounded.
  3. Add deterministic tests for new behavior.
  4. Run the local gate before opening a pull request.
  5. Never include credentials, cookies, private lead lists, or live output containing personal data.

Please use the issue tracker for design questions and bug reports. For security concerns, follow SECURITY.md instead of opening a public issue.

Roadmap

  • Add more fixture-backed provider contract tests.
  • Add configurable retention and redaction hooks for downstream persistence.
  • Improve structured company-profile extraction without turning inferred fields into facts.
  • Publish versioned container images after a maintainer review.

License and third-party software

The original orchestration code is released under the MIT License. Third-party dependencies retain their own licenses. Using this project does not override LinkedIn terms, privacy obligations, provider contracts, or organizational approval requirements.

About

Read-only, provenance-first research service for public LinkedIn company pages

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages