From 2d465affd3b7495425557425658dccc4e3c6becb Mon Sep 17 00:00:00 2001 From: PD Hall <20580126+pdhall99@users.noreply.github.com> Date: Mon, 14 Sep 2026 22:28:26 +0100 Subject: [PATCH 1/2] Fix CHANGELOG link Correct package docstring --- docs/CHANGELOG.md | 2 +- src/spanmark/__init__.py | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/CHANGELOG.md b/docs/CHANGELOG.md index ed24ec6..74b8814 100644 --- a/docs/CHANGELOG.md +++ b/docs/CHANGELOG.md @@ -15,5 +15,5 @@ - Initial public release [Unreleased]: https://github.com/pdhall99/spanmark/compare/0.1.1...HEAD -[0.1.1]: https://github.com/pdhall99/spanmark/releases/compare/0.1.0...0.1.1 +[0.1.1]: https://github.com/pdhall99/spanmark/compare/0.1.0...0.1.1 [0.1.0]: https://github.com/pdhall99/spanmark/releases/tag/0.1.0 diff --git a/src/spanmark/__init__.py b/src/spanmark/__init__.py index bebcc4c..b9ca1f0 100644 --- a/src/spanmark/__init__.py +++ b/src/spanmark/__init__.py @@ -1,4 +1,4 @@ -"""spanmark: a span annotation widget.""" +"""A lightweight span-annotation widget for interactive Python environments.""" from importlib.metadata import PackageNotFoundError, version From 98e504090f7c99b616c69c5e9f797d6dd55abb8c Mon Sep 17 00:00:00 2001 From: PD Hall <20580126+pdhall99@users.noreply.github.com> Date: Mon, 14 Sep 2026 22:32:39 +0100 Subject: [PATCH 2/2] Move use cases -> background in README --- README.md | 26 +++++++++++++------------- 1 file changed, 13 insertions(+), 13 deletions(-) diff --git a/README.md b/README.md index 1f3c901..d6fb5e0 100644 --- a/README.md +++ b/README.md @@ -8,15 +8,27 @@ Add span annotations to text datasets in Jupyter Notebook, JupyterLab, VS Code, ## Table of contents +- [Background](#background) - [Installation](#installation) - [Usage](#usage) -- [Use cases](#use-cases) - [Related tools](#related-tools) - [Compatibility and versioning](#compatibility-and-versioning) - [Acknowledgements](#acknowledgements) - [Contributing](#contributing) - [License](#license) +## Background + +spanmark helps you to make a human-verified dataset of labelled, character-based spans over document text. +Such datasets are needed for the training and evaluation of **span recognition** tasks such as + +- **Named entity recognition (NER)** — people, organizations, locations, products, dates, and other entity mentions +- **PII and sensitive-data annotation** — names, addresses, account identifiers, phone numbers, email addresses, and spans for redaction datasets +- **Keyphrase, terminology, and concept extraction** — domain terms in technical, legal, biomedical, financial, or product text +- **Slot and field extraction** — destinations, dates, quantities, order numbers, product names, and similar values in conversational or transactional text +- **Event-trigger and mention detection** — the exact text that expresses an event or concept +- **Model correction and human-in-the-loop review** — preload model suggestions, then accept, remove, relabel, or supplement them + ## Installation `spanmark` requires Python 3.10 or later (see [compatibility and versioning](#compatibility-and-versioning)). @@ -344,18 +356,6 @@ session.save() Normal session close and workflow completion also checkpoint. If the process is interrupted before that happens, keep the hidden autosave file beside the JSONL so spanmark can recover it on the next open. -## Use cases - -spanmark helps you to make a human-verified dataset of labelled, character-based spans over document text. -Such datasets are needed for the training and evaluation of **span recognition** tasks such as - -- **Named entity recognition (NER)** — people, organizations, locations, products, dates, and other entity mentions -- **PII and sensitive-data annotation** — names, addresses, account identifiers, phone numbers, email addresses, and spans for redaction datasets -- **Keyphrase, terminology, and concept extraction** — domain terms in technical, legal, biomedical, financial, or product text -- **Slot and field extraction** — destinations, dates, quantities, order numbers, product names, and similar values in conversational or transactional text -- **Event-trigger and mention detection** — the exact text that expresses an event or concept -- **Model correction and human-in-the-loop review** — preload model suggestions, then accept, remove, relabel, or supplement them - ## Related tools - [Label Studio](https://labelstud.io/) is a general-purpose data labeling across text and other modalities