Skip to content

Defining Operational Throughput: A Guide to Task Batching Limits in NumDetect #56

Description

@aiagentchat

Defining Operational Throughput: A Guide to Task Batching Limits in NumDetect

When integrating asynchronous bulk processing into your data pipeline, the efficiency of your workflow is dictated by how you structure your input files. For developers handling large-scale datasets—such as a list of 1,000,000 phone records—the strategy for batching these numbers directly impacts the success of your API submissions. You can find comprehensive details on these requirements at https://numdetect.com/api-docs.

Understanding Task Constraints

NumDetect operates as an asynchronous bulk workflow. When submitting data, it is critical to adhere to the defined per-task boundaries. According to the current API documentation, each task must contain between 500 and 500,000 phone numbers.

  • Minimum Threshold: Files containing fewer than 500 numbers are rejected upon submission and do not incur costs.
  • Maximum Threshold: A single task is capped at 500,000 records.

Decision Record: Scaling to 1,000,000 Records

If you have a dataset of 1,000,000 records, you face a choice between submitting two large tasks of 500,000 records or breaking the data into smaller, more granular chunks.

The Recommended Approach:
For a dataset of 1,000,000 records, the most efficient path is to submit two distinct tasks of 500,000 records each.

Consequences and Boundaries:

  1. Alignment with Limits: By targeting the maximum allowed batch size, you minimize the number of API calls required to process your full dataset, reducing the overhead of managing task states.
  2. File Format Requirements: Ensure your input is formatted as a TXT or CSV file with exactly one E.164-compliant phone number per line. XLSX or other spreadsheet formats are not supported.
  3. Regional Scope: Each task must be associated with one specific ISO country or region code. If your dataset spans multiple regions, you must partition your files by region before submitting them.
  4. Asynchronous Lifecycle: Remember that this is a non-real-time workflow. After submission, you must monitor the task status. Implement a configurable and non-aggressive polling policy to check for the final state of your tasks.

Implementation Checklist

  • Verify E.164 Compliance: Ensure all numbers are in the international format (beginning with a country code, maximum 15 digits).
  • Data Minimization: Under GDPR Article 5(1)(c), ensure the data provided is limited to what is necessary for your specific processing purpose.
  • Format Validation: Confirm your file is in plain TXT or CSV format.
  • Regional Grouping: Separate your records by ISO country code, as each task accepts only one region.
  • Batch Sizing: Verify that each file contains at least 500 and no more than 500,000 entries.

Takeaway

Optimizing your throughput in NumDetect is a matter of aligning your batch size with the documented 500–500,000 record constraint. By grouping your 1,000,000 records into two maximum-capacity tasks, you maintain operational efficiency while staying strictly within the supported API boundaries. For more information, visit https://numdetect.com.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions