Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 13 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,15 +8,27 @@ Add span annotations to text datasets in Jupyter Notebook, JupyterLab, VS Code,

## Table of contents

- [Background](#background)
- [Installation](#installation)
- [Usage](#usage)
- [Use cases](#use-cases)
- [Related tools](#related-tools)
- [Compatibility and versioning](#compatibility-and-versioning)
- [Acknowledgements](#acknowledgements)
- [Contributing](#contributing)
- [License](#license)

## Background

spanmark helps you to make a human-verified dataset of labelled, character-based spans over document text.
Such datasets are needed for the training and evaluation of **span recognition** tasks such as

- **Named entity recognition (NER)** — people, organizations, locations, products, dates, and other entity mentions
- **PII and sensitive-data annotation** — names, addresses, account identifiers, phone numbers, email addresses, and spans for redaction datasets
- **Keyphrase, terminology, and concept extraction** — domain terms in technical, legal, biomedical, financial, or product text
- **Slot and field extraction** — destinations, dates, quantities, order numbers, product names, and similar values in conversational or transactional text
- **Event-trigger and mention detection** — the exact text that expresses an event or concept
- **Model correction and human-in-the-loop review** — preload model suggestions, then accept, remove, relabel, or supplement them

## Installation

`spanmark` requires Python 3.10 or later (see [compatibility and versioning](#compatibility-and-versioning)).
Expand Down Expand Up @@ -344,18 +356,6 @@ session.save()
Normal session close and workflow completion also checkpoint.
If the process is interrupted before that happens, keep the hidden autosave file beside the JSONL so spanmark can recover it on the next open.

## Use cases

spanmark helps you to make a human-verified dataset of labelled, character-based spans over document text.
Such datasets are needed for the training and evaluation of **span recognition** tasks such as

- **Named entity recognition (NER)** — people, organizations, locations, products, dates, and other entity mentions
- **PII and sensitive-data annotation** — names, addresses, account identifiers, phone numbers, email addresses, and spans for redaction datasets
- **Keyphrase, terminology, and concept extraction** — domain terms in technical, legal, biomedical, financial, or product text
- **Slot and field extraction** — destinations, dates, quantities, order numbers, product names, and similar values in conversational or transactional text
- **Event-trigger and mention detection** — the exact text that expresses an event or concept
- **Model correction and human-in-the-loop review** — preload model suggestions, then accept, remove, relabel, or supplement them

## Related tools

- [Label Studio](https://labelstud.io/) is a general-purpose data labeling across text and other modalities
Expand Down
2 changes: 1 addition & 1 deletion docs/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,5 +15,5 @@
- Initial public release

[Unreleased]: https://github.com/pdhall99/spanmark/compare/0.1.1...HEAD
[0.1.1]: https://github.com/pdhall99/spanmark/releases/compare/0.1.0...0.1.1
[0.1.1]: https://github.com/pdhall99/spanmark/compare/0.1.0...0.1.1
[0.1.0]: https://github.com/pdhall99/spanmark/releases/tag/0.1.0
2 changes: 1 addition & 1 deletion src/spanmark/__init__.py
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
"""spanmark: a span annotation widget."""
"""A lightweight span-annotation widget for interactive Python environments."""

from importlib.metadata import PackageNotFoundError, version

Expand Down