POST /v1/nlp/tokenize
Tokeniser
Sentence and word boundaries done properly, including abbreviations and decimals.
Split text into words or sentences now No key, no code — 3 credits either way.
curl -X POST "$DOATHING_API/v1/nlp/tokenize" \
-H "x-api-key: $DOATHING_KEY" \
-H "content-type: application/json" \
-d '{"text": "Dr. Smith arrived at 3 p.m. He was late."}'
Sentence and word boundaries done properly, including abbreviations and decimals.
Split text into sentences or word tokens with character offsets, using NLTK's Punkt tokeniser.
Input
Send either a text string or a file object. A file is decoded as UTF-8 and must be one of text/plain, text/markdown, text/html, text/csv.
Response
The result carries your remaining balance alongside it, so you can track spend without a second call.
{
"unit": "sentences",
"count": 2,
"sentences": [
"Dr. Smith arrived at 3 p.m.",
"He was late."
],
"request_id": "37f01edb-0163-42a1-ac51-0acaef979800",
"credits_remaining": 96
}
Cost
3 credits per call, whether it is run from the site or from the API — the credential differs, the price does not. A new account starts with 20 credits.
A rejected request still costs a credit: the authorizer decrements before the tool validates. A call rejected for a missing or invalid key is free.
Parameters
Generated from the endpoint’s own validation, so this is exactly what it accepts. A body field goes at the top level; an option goes inside options.
| Name | In | Type | Default | Notes |
|---|---|---|---|---|
| unit | options | string | "words" | words or sentences. one of sentences, words |
| strip_punctuation | options | boolean | false | Drop punctuation tokens. |