diff --git a/README.md b/README.md index 753d7a1..89e05a8 100644 --- a/README.md +++ b/README.md @@ -7,12 +7,20 @@ replacement for the `json` module's `loads`, `load`, `dumps` and `dump`, and it can be over 3 times faster than the standard `json.loads` and `json.dumps`. When you only need part of a document, its lazy `parse` function is faster still. It also reads streams of documents (NDJSON, JSON -Lines). +Lines) and writes JSON as `bytes` (`dumpb`). + +With pip: ```sh pip install fastsimdjson ``` +With uv: + +```sh +uv pip install fastsimdjson # in a uv project: uv add fastsimdjson +``` + Wheels are available for Linux, macOS and Windows, for Python 3.10 to 3.14, including free-threaded Python 3.14. @@ -68,12 +76,12 @@ doc["search_metadata"].as_dict() # convert a subtree, like loads `json.loads` accepts (an overflowing number, an unpaired surrogate) is returned as plain Python objects, as `loads` would return it. -### `dumps` and `dump`: writing JSON +### `dumps`, `dumpb` and `dump`: writing JSON ```python fastsimdjson.dumps({"a": [1, 2.5, None]}) # '{"a": [1, 2.5, null]}' fastsimdjson.dumps(obj, indent=2, sort_keys=True) -fastsimdjson.dumps(obj, separators=(",", ":"), ensure_ascii=False) +fastsimdjson.dumpb(obj, separators=(",", ":"), ensure_ascii=False) # bytes with open("out.json", "w", encoding="utf-8") as f: fastsimdjson.dump(obj, f) ``` @@ -85,7 +93,10 @@ escapes and the default separators. `ensure_ascii`, `indent`, `separators`, passed to `json.dumps` itself, which produces the result or raises its usual exception: a `cls` argument or other encoder options, `skipkeys`, a circular reference, `NaN` with `allow_nan=False`, a key or a value that `json` cannot -serialize. `dump(obj, fp, **kw)` writes `dumps(obj, **kw)` to `fp`. +serialize. `dumpb(obj, **kw)` returns the same text as UTF-8 `bytes`, exactly +`json.dumps(obj, **kw).encode()`, without building a `str` first: use it to +write to a binary file or a socket. `dump(obj, fp, **kw)` writes +`dumps(obj, **kw)` to `fp`. ### Files @@ -118,16 +129,15 @@ separated: | `"whitespace"` (default) | documents separated by white space, including NDJSON and JSON Lines | | `"lines"` | one document per line (NDJSON, JSON Lines) | | `"json_seq"` | RFC 7464 JSON text sequences (each document preceded by `\x1e`) | -| `"comma"` | documents separated by commas: `{...}, {...}` | +| `"comma"` | documents separated by commas: `{...}, {...}` (simdjson also accepts white space between them) | | `"array"` | the elements of one array: `[{...}, {...}]` | simdjson parses the input in batches (`batch_size`, 1 MB by default); a -larger document is handled automatically. With `"whitespace"` and `"lines"`, -documents that simdjson rejects are handled as in `loads`: a document that -`json` accepts is returned, otherwise `JSONDecodeError` reports `json`'s -message and the position in the whole input. A truncated last document is an -error. The views returned by `parse_many` remain valid after the iterator -moves on. +larger document is handled automatically. In every format, documents that +simdjson rejects are handled as in `loads`: a document that `json` accepts is +returned, otherwise `JSONDecodeError` reports `json`'s message and the +position in the whole input. A truncated last document is an error. The +views returned by `parse_many` remain valid after the iterator moves on. ### `release` @@ -182,8 +192,8 @@ why `PYTHONPATH=src` is required for the in-place build. `tests/test_loads.py` compares `loads` with `json.loads` on types and key order; `tests/test_lazy.py` checks the views returned by `parse` the same way, -`tests/test_dumps.py` compares `dumps` with `json.dumps` (output and -exceptions), and `tests/test_stream.py` and `tests/test_files.py` cover +`tests/test_dumps.py` compares `dumps` and `dumpb` with `json.dumps` (output +and exceptions), and `tests/test_stream.py` and `tests/test_files.py` cover streams and files. It covers scalars, integers past 64 bits, UTF-8 strings at every length from 0 to 199, the key cache, random documents, rejected input, deep nesting, padding at a page boundary, a saturated array count, reference @@ -199,26 +209,38 @@ pytest tests The suite builds an ~80 MB document and a list of 16,777,221 integers, so give it some RAM. -To time `loads` against `json.loads` and orjson on those files -(`bench_lazy.py` times `parse` against pysimdjson and cysimdjson, -`bench_dumps.py` times `dumps`, and `bench_many.py` times `loads_many`): +The benchmark scripts need simdjson-data and the `bench` extra (orjson, +msgspec, pysimdjson, cysimdjson). `bench.py` times `loads` against +`json.loads` and orjson, `bench_lazy.py` times `parse` against pysimdjson and +cysimdjson, `bench_dumps.py` times `dumps` and `dumpb`, and `bench_many.py` +times `loads_many`. + +pip: ```sh python -m pip install -e ".[bench]" python bench.py +python bench_lazy.py +python bench_dumps.py +python bench_many.py ``` +uv: + ```sh uv pip install -e ".[bench]" python bench.py +python bench_lazy.py +python bench_dumps.py +python bench_many.py ``` ## Benchmarks Intel Xeon Gold 6548N (Emerald Rapids), one core, Python 3.14.6, -fastsimdjson 0.2.0, the 22 files of -[simdjson-data](https://github.com/simdjson/simdjson-data). The scripts and -the full results are in +fastsimdjson 0.3.0 (development version, simdjson 5.0.2), the 22 files of +[simdjson-data](https://github.com/simdjson/simdjson-data). The scripts are in +this repository and in [the blog repository](https://github.com/lemire/Code-used-on-Daniel-Lemire-s-blog/tree/master/2026/09/pysimdjson). ### Whole documents @@ -229,24 +251,25 @@ geometric mean over the 22 files (higher is better). | parser | GB/s | vs `json.loads` | |---|---:|---:| | json (standard library) | 0.22 | 1.00× | -| simplejson 4.1.2 | 0.23 | 1.05× | +| simplejson 4.2.0 | 0.23 | 1.05× | | python-rapidjson 1.25 | 0.24 | 1.10× | -| ujson 6.0.0 | 0.36 | 1.65× | -| cysimdjson 26.27 | 0.43 | 1.94× | +| ujson 6.0.0 | 0.37 | 1.66× | +| cysimdjson 26.27 | 0.43 | 1.96× | | pysimdjson 7.0.2 | 0.44 | 1.98× | -| msgspec 0.22.0 | 0.53 | 2.41× | -| orjson 3.12.0 | 0.60 | 2.73× | -| **fastsimdjson `loads`** | **0.77** | **3.49×** | +| msgspec 0.22.0 | 0.53 | 2.42× | +| orjson 3.12.0 | 0.60 | 2.74× | +| **fastsimdjson `loads`** | **0.78** | **3.53×** | fastsimdjson is the fastest on 21 of the 22 files; orjson is slightly faster on `numbers.json`, an array of floating-point numbers. Part of the gain comes from pausing the garbage collector while the objects are built: if the collector is disabled for every parser, fastsimdjson's lead over orjson drops -from 1.28× to 1.18×. yyjson 4.0.6 is left out: it returns wrong strings for -non-ASCII text. +from 1.29× to 1.17×, and orjson is slightly faster on `canada.json`, +`mesh.json` and `numbers.json`. yyjson 4.0.6 is left out: it returns wrong +strings for non-ASCII text. Parsing is no longer the bottleneck. simdjson alone parses these files at -3.0 GB/s. It accounts for about a third of the time of `loads`; the rest goes +3.1 GB/s. It accounts for about a third of the time of `loads`; the rest goes into creating Python objects. Freeing those objects later costs about a sixth of the total. Even if parsing took no time at all, `loads` would be less than 1.5 times faster. @@ -258,90 +281,72 @@ and the screen name of the 100 statuses of `twitter.json`: | method | µs | |---|---:| -| `json.loads` | 3879 | -| orjson | 1008 | -| fastsimdjson `loads` | 860 | +| `json.loads` | 3922 | +| orjson | 1009 | +| fastsimdjson `loads` | 861 | | msgspec (typed `Struct`) | 336 | -| cysimdjson (lazy) | 235 | -| pysimdjson (lazy) | 183 | -| **fastsimdjson `parse`** | **155** | - -Here `parse` is 25 times faster than `json.loads` and 5.5 times faster than -`loads`. Most of its time is the simdjson parse itself: reading the 200 -values takes less than 20 µs. Compared with pysimdjson on other tasks -(µs, lower is better): - -| file | task | fastsimdjson `parse` | fastsimdjson `loads` | pysimdjson | -|---|---|---:|---:|---:| -| twitter | open | 140 | 676 | 156 | -| citm_catalog | open | 369 | 1612 | 472 | -| citm_catalog | extract | 394 | 2142 | 504 | -| gsoc-2018 | open | 615 | 2206 | 813 | -| twitter | visit all | 1985 | 2183 | 3033 | -| canada | visit all | 17528 | 18998 | 19750 | -| twitter_api_response | open | 3.6 | 13.6 | 3.4 | - -"open" parses the document and looks at its root; "extract" collects the -start time of every performance; "visit all" walks every value through the -views (with `loads`: through the dict). When you visit everything, `parse` -is about as fast as `loads`. On very small documents (15 KB), pysimdjson's -`parse` is marginally faster. - -### Writing JSON with `dumps` - -Same machine, fastsimdjson 0.3.0, the objects of the 22 files. With the -default arguments, `dumps` returns exactly what `json.dumps` returns and is -3.4 times faster (geometric mean; from 2.5 times on text-heavy files to 8 -times on files full of numbers). Microseconds: - -| file | `json.dumps` | fastsimdjson `dumps` | orjson | msgspec | +| cysimdjson (lazy) | 229 | +| pysimdjson (lazy) | 179 | +| **fastsimdjson `parse`** | **158** | + +Here `parse` is about 25 times faster than `json.loads` and 5.5 times faster +than `loads`. Most of its time is the simdjson parse itself. + +### Writing JSON with `dumps` and `dumpb` + +With the default arguments, `dumps` returns exactly the `str` of `json.dumps` +and is 3.3 times faster (geometric mean over the 22 files; from 2.3 times on +text-heavy files to 8 times on files full of numbers). + +orjson and msgspec produce compact UTF-8 `bytes`: no spaces after separators, +non-ASCII characters left as they are. The fair comparison is with the same +output: `dumpb(obj, separators=(",", ":"), ensure_ascii=False)`, and +`json.dumps` with the same arguments followed by `.encode()`. These four +produce identical bytes on 18 of the 22 files; on the others they differ only +in how some floats are written (`1e-05` as in Python's `repr`, against +`0.00001`). Microseconds: + +| file | `json.dumps(...).encode()` | fastsimdjson `dumpb` | orjson | msgspec | |---|---:|---:|---:|---:| -| twitter | 1482 | 529 | 199 | 347 | -| citm_catalog | 2702 | 1071 | 427 | 494 | -| github_events | 158 | 43 | 19 | 30 | -| canada | 38622 | 4843 | 2923 | 3653 | -| numbers | 2515 | 349 | 198 | 332 | - -orjson and msgspec are faster still, but they produce something else: they -return `bytes`, without spaces after separators and without escaping -non-ASCII characters. Compared with `separators=(",", ":")` and -`ensure_ascii=False`, the closest `dumps` settings, orjson is 2.9 times -faster and msgspec 1.8 times faster (geometric means). +| twitter | 1798 | 549 | 200 | 348 | +| citm_catalog | 2956 | 1115 | 430 | 493 | +| github_events | 173 | 45 | 18 | 30 | +| gsoc-2018 | 15309 | 1848 | 546 | 1423 | +| canada | 38095 | 4708 | 2921 | 3636 | +| numbers | 2515 | 352 | 197 | 333 | + +`dumpb` is 3.8 times faster than `json.dumps(...).encode()` (geometric mean; +2.4 to 8.3 times). orjson is faster still, by 2.6 times, and msgspec by 1.6 +times (geometric means). Unlike them, `dumps` and `dumpb` accept every argument +of `json.dumps` and produce its exact output. ### Streams with `loads_many` 20 MB of NDJSON (5268 objects and arrays, one per line), made from the same -files: +files; best of three runs: | method | ms | GB/s | |---|---:|---:| -| `json.loads` on each line | 167 | 0.12 | -| orjson on each line | 81 | 0.24 | -| fastsimdjson `loads` on each line | 61 | 0.32 | +| `json.loads` on each line | 169 | 0.12 | +| orjson on each line | 77 | 0.26 | +| msgspec `decode` on each line | 74 | 0.27 | +| fastsimdjson `loads` on each line | 60 | 0.33 | +| msgspec `decode_lines` | 53 | 0.37 | | fastsimdjson `loads_many` | 52 | 0.38 | -| fastsimdjson `parse_many` (views only) | 20 | 0.96 | +| fastsimdjson `parse_many` (views only) | 21 | 0.94 | + +msgspec's `decode_lines` and `loads_many` are on par (`loads_many` is about 2% +faster). `parse_many` only creates the views; reading values from them adds to +its time. ## Limitations -* `dumps` returns a `str`, like `json.dumps`; there is no option to return - `bytes`. -* In streams, a document that is a bare number (`3.14` alone on its line) - is slow to parse: simdjson copies the rest of the batch for each one. Streams - of objects and arrays are not affected. -* With the `"json_seq"`, `"comma"` and `"array"` stream formats, documents - that simdjson rejects raise `JSONDecodeError` with simdjson's message; - there is no fallback to `json`. * The simdjson parser and the key and string caches are thread-local. `release()` frees the parser and the cached strings retained by the calling thread. A parser that grows past 64 MB is freed on its own at the end of that call; its caches stay. A thread that exits without `release()` leaves - its cached strings behind. The module is marked free-threading compatible - (`Py_MOD_GIL_NOT_USED` on Python 3.13 and newer), so importing it on a - free-threaded build does not re-enable the GIL. Subinterpreters are not - supported. On a free-threaded build, `bytearray` and `memoryview` inputs - are copied before parsing. -* A document simdjson rejects is reparsed with `json.loads`. `loads` returns - that value when `json.loads` accepts it, which is how overflow to infinity - is handled. An exception is raised only when `json.loads` also fails. - `JSONDecodeError` is re-raised as `fastsimdjson.JSONDecodeError` with the - same message, document, and position. Any other exception propagates. + its cached strings behind. +* The module is marked free-threading compatible (`Py_MOD_GIL_NOT_USED` on + Python 3.13 and newer), so importing it on a free-threaded build does not + re-enable the GIL. On a free-threaded build, `bytearray` and `memoryview` + inputs are copied before parsing. Subinterpreters are not supported. diff --git a/bench_dumps.py b/bench_dumps.py index 31dc331..608adf9 100644 --- a/bench_dumps.py +++ b/bench_dumps.py @@ -1,9 +1,12 @@ -"""Time dumps against json.dumps, orjson and msgspec on simdjson-data. +"""Time dumps and dumpb against json.dumps, orjson and msgspec on simdjson-data. -Two settings: - default: json.dumps(obj) and fastsimdjson.dumps(obj) (identical str output); - compact: separators=(",", ":"), ensure_ascii=False, the closest to what - orjson.dumps and msgspec.json.encode produce (they return bytes). +Two comparisons: + str: json.dumps(obj) against fastsimdjson.dumps(obj), the same str; + bytes: UTF-8 bytes in the compact form that orjson and msgspec produce + (no spaces, non-ASCII characters not escaped): + json.dumps(obj, separators=(",", ":"), ensure_ascii=False).encode(), + fastsimdjson.dumpb(obj, separators=(",", ":"), ensure_ascii=False), + orjson.dumps(obj) and msgspec.json.encode(obj). """ import glob import json @@ -46,24 +49,24 @@ def best(fn, obj, target=0.02, rounds=11): def main(): - cols = ["json", "fast", "json compact", "fast compact"] + cols = ["json str", "fast str", "json bytes", "fast bytes"] if orjson: cols.append("orjson") if msgspec: cols.append("msgspec") - print(f"{'file':36s} " + " ".join(f"{c:>13s}" for c in cols) + " (us)") + print(f"{'file':36s} " + " ".join(f"{c:>11s}" for c in cols) + " (us)") for f in FILES: obj = json.loads(open(f, "rb").read()) assert fastsimdjson.dumps(obj) == json.dumps(obj) - assert fastsimdjson.dumps(obj, **COMPACT) == json.dumps(obj, **COMPACT) + assert fastsimdjson.dumpb(obj, **COMPACT) == json.dumps(obj, **COMPACT).encode() row = [best(json.dumps, obj), best(fastsimdjson.dumps, obj), - best(lambda o: json.dumps(o, **COMPACT), obj), - best(lambda o: fastsimdjson.dumps(o, **COMPACT), obj)] + best(lambda o: json.dumps(o, **COMPACT).encode(), obj), + best(lambda o: fastsimdjson.dumpb(o, **COMPACT), obj)] if orjson: row.append(best(orjson.dumps, obj)) if msgspec: row.append(best(msgspec.json.encode, obj)) - print(f"{os.path.basename(f):36s} " + " ".join(f"{t:13.1f}" for t in row)) + print(f"{os.path.basename(f):36s} " + " ".join(f"{t:11.1f}" for t in row)) if __name__ == "__main__": diff --git a/bench_many.py b/bench_many.py index fb5fc6e..a5259bc 100644 --- a/bench_many.py +++ b/bench_many.py @@ -16,6 +16,10 @@ import orjson except ImportError: orjson = None +try: + import msgspec +except ImportError: + msgspec = None DATA = os.environ.get("JSONDIR", "simdjson-data/jsonexamples") @@ -55,6 +59,10 @@ def main(): } if orjson: methods["orjson.loads per line"] = lambda d: [orjson.loads(l) for l in d.splitlines() if l] + if msgspec: + decoder = msgspec.json.Decoder() + methods["msgspec decode per line"] = lambda d: [decoder.decode(l) for l in d.splitlines() if l] + methods["msgspec decode_lines"] = decoder.decode_lines ref = methods["json.loads per line"](data) for name, fn in methods.items(): if "parse_many" not in name: diff --git a/pyproject.toml b/pyproject.toml index cedaedb..062c188 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -34,7 +34,7 @@ Issues = "https://github.com/simdjson/fastpysimdjson/issues" [project.optional-dependencies] test = ["pytest"] -bench = ["orjson"] +bench = ["orjson", "msgspec", "pysimdjson", "cysimdjson"] [tool.cibuildwheel] skip = ["pp*"] diff --git a/src/fastsimdjson.cpp b/src/fastsimdjson.cpp index 24717c4..ef548ef 100644 --- a/src/fastsimdjson.cpp +++ b/src/fastsimdjson.cpp @@ -1863,11 +1863,24 @@ bool configure(Encoder &e, PyObject *kwargs, bool *error) { return true; } -PyObject *dumps(PyObject *, PyObject *args, PyObject *kwargs) { +// json.dumps(*args, **kwargs), as a str, or encoded as UTF-8 bytes. +PyObject *call_json_dumps(PyObject *args, PyObject *kwargs, bool as_bytes) { + PyObject *s = PyObject_Call(json_dumps, args, kwargs); + if (s == nullptr || !as_bytes) { + return s; + } + PyObject *b = PyUnicode_AsUTF8String(s); + Py_DECREF(s); + return b; +} + +// dumps returns the str of json.dumps; dumpb returns it encoded as UTF-8. +template +PyObject *serialize(PyObject *, PyObject *args, PyObject *kwargs) { Encoder e; bool error = false; if (PyTuple_GET_SIZE(args) != 1 || !configure(e, kwargs, &error)) { - return error ? nullptr : PyObject_Call(json_dumps, args, kwargs); + return error ? nullptr : call_json_dumps(args, kwargs, AsBytes); } int rc; try { @@ -1879,12 +1892,23 @@ PyObject *dumps(PyObject *, PyObject *args, PyObject *kwargs) { return nullptr; } if (rc == ENCODE_FALLBACK) { - return PyObject_Call(json_dumps, args, kwargs); + return call_json_dumps(args, kwargs, AsBytes); + } + if (AsBytes) { // the output is UTF-8 (ASCII with ensure_ascii) + return PyBytes_FromStringAndSize(e.out.data(), Py_ssize_t(e.out.size())); } return e.ascii_output ? new_ascii(e.out.data(), e.out.size()) : PyUnicode_DecodeUTF8(e.out.data(), Py_ssize_t(e.out.size()), nullptr); } +PyObject *dumps(PyObject *self, PyObject *args, PyObject *kwargs) { + return serialize(self, args, kwargs); +} + +PyObject *dumpb(PyObject *self, PyObject *args, PyObject *kwargs) { + return serialize(self, args, kwargs); +} + // dump(obj, fp, **kw): fp.write(dumps(obj, **kw)). PyObject *dump(PyObject *, PyObject *args, PyObject *kwargs) { if (PyTuple_GET_SIZE(args) != 2) { @@ -2067,26 +2091,31 @@ bool stream_start(StreamObject *s, size_t at) { PyErr_NoMemory(); return false; } - // parse_many does not pass number_as_string on to the parser - // implementation (simdjson 5.0.2), so big integers would fail: create the - // implementation now and set the flag on it; later reallocations keep it. - if (!s->p->implementation) { - simdjson::error_code alloc_err = s->p->allocate(s->batch_size); - if (alloc_err) { - stream_error(s, simdjson::error_message(alloc_err), at); - return false; + // The first start parses the whole input in its format. A restart (after a + // document that json decoded, or with larger batches) begins after a + // document: comma-separated documents restart at the next one, and the + // rest of an array is parsed as comma-separated documents before its ']'. + using simdjson::stream_format; + stream_format fmt = s->format; + size_t end = s->len; + if (at > 0) { + if (fmt == stream_format::comma_delimited_array) { + fmt = stream_format::comma_delimited; + end = s->close < s->len ? s->close : s->len; + } + if (fmt == stream_format::comma_delimited) { + at = skip_separators(s, at); } } - s->p->implementation->_number_as_string = true; auto r = s->p->parse_many(reinterpret_cast(s->buf) + at, - s->len - at, s->batch_size, s->format); + at < end ? end - at : 0, s->batch_size, fmt); simdjson::error_code err = std::move(r).get(*s->stream); if (err) { stream_error(s, simdjson::error_message(err), at); return false; } s->it = s->stream->begin(); - // Restarts happen after the prefix: offsets are relative to it only at 0. + // At 0, offsets are relative to what follows the prefix. s->base = at == 0 ? s->prefix : at; s->active = true; return true; @@ -2132,16 +2161,10 @@ PyObject *stream_value(StreamObject *s) { return d == nullptr ? PyErr_NoMemory() : view_of(d); } -// simdjson rejected the document after last_end. For whitespace-separated -// documents, json decides, as in loads: return the value it accepts and -// resume after it, or raise its error. -PyObject *stream_fallback(StreamObject *s, simdjson::error_code err) { +// simdjson rejected the document after last_end. json decides, as in loads: +// return the value it accepts and resume after it, or raise its error. +PyObject *stream_fallback(StreamObject *s) { size_t at = skip_separators(s, s->last_end); - if (s->format != simdjson::stream_format::whitespace_delimited && - s->format != simdjson::stream_format::newline_delimited) { - stream_error(s, simdjson::error_message(err), at); - return nullptr; - } PyObject *text = stream_text(s); if (text == nullptr) { return nullptr; @@ -2160,6 +2183,23 @@ PyObject *stream_fallback(StreamObject *s, simdjson::error_code err) { return nullptr; } s->last_end = byte_offset(s->buf, s->len, at, end - start); + // Between documents, the formats other than white space need their + // separator: json decoded one document, not the separator after it. + using simdjson::stream_format; + size_t next = s->last_end; + while (next < s->len && json_space(s->buf[next])) { + next++; + } + char sep = s->format == stream_format::json_sequence ? '\x1e' + : s->format == stream_format::comma_delimited || + s->format == stream_format::comma_delimited_array + ? ',' + : 0; + if (sep != 0 && next < s->len && next != s->close && s->buf[next] != sep) { + Py_DECREF(value); + stream_error(s, sep == ',' ? "Expecting ',' delimiter" : "Expecting record separator", next); + return nullptr; + } s->active = false; // resume with simdjson after this document // As with parse(), a document that only json accepts is returned as plain // Python objects. @@ -2178,12 +2218,9 @@ PyObject *stream_next_locked(StreamObject *s) { } } if (!(s->it != s->stream->end())) { - // truncated_bytes() is not reliable for every format: look at what - // follows the last document instead. An incomplete document at the - // end: let json report it. - return skip_separators(s, s->last_end) < s->len - ? stream_fallback(s, simdjson::TAPE_ERROR) - : nullptr; + // Anything but separators after the last document is an incomplete + // document: let json report it. + return skip_separators(s, s->last_end) < s->len ? stream_fallback(s) : nullptr; } simdjson::dom::element el; simdjson::error_code err = (*s->it).get(el); @@ -2202,7 +2239,7 @@ PyObject *stream_next_locked(StreamObject *s) { s->active = false; continue; } - return stream_fallback(s, err); + return stream_fallback(s); } } @@ -2407,6 +2444,9 @@ PyMethodDef methods[] = { {"dumps", reinterpret_cast(reinterpret_cast(dumps)), METH_VARARGS | METH_KEYWORDS, "Serialize obj to a JSON str. Same arguments and output as json.dumps."}, + {"dumpb", reinterpret_cast(reinterpret_cast(dumpb)), + METH_VARARGS | METH_KEYWORDS, + "Serialize obj to JSON as UTF-8 bytes: json.dumps(obj, **kw).encode()."}, {"dump", reinterpret_cast(reinterpret_cast(dump)), METH_VARARGS | METH_KEYWORDS, "Serialize obj as JSON to fp (a file with a write method), like json.dump."}, diff --git a/tests/test_dumps.py b/tests/test_dumps.py index 31e4082..d5b67a3 100644 --- a/tests/test_dumps.py +++ b/tests/test_dumps.py @@ -250,3 +250,51 @@ def test_dump(): assert a.getvalue() == b.getvalue() with pytest.raises(TypeError): fastsimdjson.dump(obj) + + +def expected_bytes(obj, **kw): + try: + return ("ok", json.dumps(obj, **kw).encode()) + except Exception as e: # noqa: BLE001 + return ("exc", type(e), str(e)) + + +def got_bytes(obj, **kw): + try: + return ("ok", fastsimdjson.dumpb(obj, **kw)) + except Exception as e: # noqa: BLE001 + return ("exc", type(e), str(e)) + + +@pytest.mark.parametrize("path", sorted(glob.glob(os.path.join(DATA, "*.json")))) +def test_dumpb_files(path): + obj = json.loads(open(path, "rb").read()) + for kw in OPTIONS: + out = fastsimdjson.dumpb(obj, **kw) + assert type(out) is bytes and out == json.dumps(obj, **kw).encode() + + +@pytest.mark.parametrize("seed", range(20)) +def test_dumpb_random(seed): + obj = random_value(random.Random(seed)) + for kw in OPTIONS: + assert fastsimdjson.dumpb(obj, **kw) == json.dumps(obj, **kw).encode() + + +def test_dumpb_fallbacks_and_errors(): + a = [1] + a.append(a) + cases = [ + ([float("nan")], {"allow_nan": False}), + ({(1, 2): 3}, {}), + ({(1, 2): 3}, {"skipkeys": True}), + ([{1, 2}], {}), + ([{1, 2}], {"default": sorted}), + (a, {}), + (["\ud800"], {}), + (["\ud800"], {"ensure_ascii": False}), # json's str cannot be UTF-8 + (["é\U0001f600"], {"ensure_ascii": False}), + ([object()], {"cls": None, "default": str}), + ] + for obj, kw in cases: + assert got_bytes(obj, **kw) == expected_bytes(obj, **kw), (obj, kw) diff --git a/tests/test_stream.py b/tests/test_stream.py index 5f5930c..7b7903b 100644 --- a/tests/test_stream.py +++ b/tests/test_stream.py @@ -193,3 +193,32 @@ def test_byte_order_mark_errors(fmt): with pytest.raises(fastsimdjson.JSONDecodeError) as info: list(fastsimdjson.loads_many(bad, format=fmt)) assert (info.value.msg, info.value.pos) == expected + + +@pytest.mark.parametrize("fmt", ["json_seq", "comma", "array"]) +def test_json_fallback_other_formats(fmt): + # As with white space: documents that only json accepts are decoded by + # json, and errors are json's, at their position in the whole input. + docs = ['{"a": 1}', '{"v": 1e400}', '["\\ud800"]', '{"c": 3}'] + expected = [{"a": 1}, {"v": float("inf")}, ["\ud800"], {"c": 3}] + join = {"json_seq": lambda d: "".join("\x1e" + x + "\n" for x in d), + "comma": lambda d: ", ".join(d), + "array": lambda d: "[" + ", ".join(d) + "]"}[fmt] + assert list(fastsimdjson.loads_many(join(docs), format=fmt)) == expected + got = [to_py(x) for x in fastsimdjson.parse_many(join(docs), format=fmt)] + assert got == expected + bad = join(['{"a": 1}', '{"v": 1e400}', '{"b": ]']) + with pytest.raises(fastsimdjson.JSONDecodeError) as info: + list(fastsimdjson.loads_many(bad, format=fmt)) + assert info.value.msg == "Expecting value" + assert bad[info.value.pos] == "]" + + +def test_fallback_needs_the_separator(): + # After json decodes a document, the next one must follow a separator. + with pytest.raises(fastsimdjson.JSONDecodeError) as info: + list(fastsimdjson.loads_many("1e400 2", format="json_seq")) + assert info.value.msg == "Expecting record separator" + with pytest.raises(fastsimdjson.JSONDecodeError) as info: + list(fastsimdjson.loads_many("[1e400 2]", format="array")) + assert info.value.msg == "Expecting ',' delimiter" diff --git a/vendor/simdjson.cpp b/vendor/simdjson.cpp index 5a263dd..81cd2fa 100644 --- a/vendor/simdjson.cpp +++ b/vendor/simdjson.cpp @@ -1,4 +1,4 @@ -/* auto-generated on 2026-10-01 09:33:54 -0400. version 5.0.2 Do not edit! */ +/* auto-generated on 2026-10-04 09:02:57 -0400. version 5.0.2 Do not edit! */ /* including simdjson.cpp: */ /* begin file simdjson.cpp */ #define SIMDJSON_SRC_SIMDJSON_CPP @@ -16222,7 +16222,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -19489,6 +19494,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -19567,7 +19574,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -19654,6 +19661,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -19691,7 +19700,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -21564,10 +21573,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -23923,7 +23935,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -27037,6 +27054,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -27115,7 +27134,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -27202,6 +27221,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -27239,7 +27260,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -29112,10 +29133,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -31447,7 +31471,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -34560,6 +34589,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -34638,7 +34669,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -34725,6 +34756,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -34762,7 +34795,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -36635,10 +36668,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -39128,7 +39164,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -42354,6 +42395,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -42432,7 +42475,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -42519,6 +42562,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -42556,7 +42601,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -44429,10 +44474,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -47159,7 +47207,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -50690,6 +50743,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -50768,7 +50823,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -50855,6 +50910,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -50892,7 +50949,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -52765,10 +52822,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -55042,7 +55102,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -58089,6 +58154,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -58167,7 +58234,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -58254,6 +58321,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -58291,7 +58360,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -60164,10 +60233,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -62378,7 +62450,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -65392,6 +65469,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -65470,7 +65549,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -65557,6 +65636,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -65594,7 +65675,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -67467,10 +67548,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -69701,7 +69785,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -73112,6 +73201,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -73190,7 +73281,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -73277,6 +73368,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -73314,7 +73407,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -75187,10 +75280,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } @@ -77005,7 +77101,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -78751,6 +78852,8 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // RS as a scalar, making the digit a scalar continuation, not a start. // We must: (1) remove RS from structural_indexes, and (2) for scalars, add the // actual value start position. + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_rs_pos = 0; uint32_t rs_count = 0; @@ -78829,7 +78932,7 @@ simdjson_inline uint32_t find_next_document_index_json_sequence( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { // Only RS markers here: the last one opens a record continuing past the @@ -78916,6 +79019,8 @@ simdjson_inline uint32_t filter_comma_delimited( // Track depth to identify root-level commas (depth 0) int depth = 0; + // The EOF sentinel: len, or where a discarded unclosed string starts. + const uint32_t sentinel = parser.structural_indexes[parser.n_structural_indexes]; uint32_t write_idx = 0; uint32_t last_root_comma_pos = 0; uint32_t root_comma_count = 0; @@ -78953,7 +79058,7 @@ simdjson_inline uint32_t filter_comma_delimited( // past the end: restore the EOF sentinel that stage 1 had planted there, which // document_stream::truncated_bytes() reads after a final batch. parser.n_structural_indexes = write_idx; - parser.structural_indexes[write_idx] = uint32_t(len); + parser.structural_indexes[write_idx] = sentinel; if (parser.n_structural_indexes == 0) { return 0; } @@ -80210,10 +80315,13 @@ simdjson_warn_unused simdjson_inline error_code tape_builder_impl::vis // practice unless you are in the strange scenario where you have many JSON // documents made of single atoms. // - std::unique_ptrcopy(new (std::nothrow) uint8_t[iter.remaining_len() + SIMDJSON_PADDING]); + // In a stream, the input goes on with other documents: copy up to the next + // structural only, not to the end of the batch. + const size_t len = (std::min)(iter.remaining_len(), size_t(*iter.next_structural) - size_t(*(iter.next_structural - 1))); + std::unique_ptrcopy(new (std::nothrow) uint8_t[len + SIMDJSON_PADDING]); if (copy.get() == nullptr) { return MEMALLOC; } - std::memcpy(copy.get(), value, iter.remaining_len()); - std::memset(copy.get() + iter.remaining_len(), ' ', SIMDJSON_PADDING); + std::memcpy(copy.get(), value, len); + std::memset(copy.get() + len, ' ', SIMDJSON_PADDING); error_code error = visit_number(iter, copy.get()); return error; } diff --git a/vendor/simdjson.h b/vendor/simdjson.h index df5bc55..0818b46 100644 --- a/vendor/simdjson.h +++ b/vendor/simdjson.h @@ -1,4 +1,4 @@ -/* auto-generated on 2026-10-01 09:33:54 -0400. version 5.0.2 Do not edit! */ +/* auto-generated on 2026-10-04 09:02:57 -0400. version 5.0.2 Do not edit! */ /* including simdjson.h: */ /* begin file simdjson.h */ #ifndef SIMDJSON_H @@ -11008,6 +11008,7 @@ inline void document_stream::start() noexcept { if (error) { return; } error = parser->ensure_capacity(batch_size); if (error) { return; } + parser->implementation->_number_as_string = parser->number_as_string(); // Always run the first stage 1 parse immediately batch_start = 0; error = run_stage1(*parser, batch_start); @@ -17430,7 +17431,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -20256,7 +20262,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -23559,7 +23570,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -26862,7 +26878,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -30280,7 +30301,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -34005,7 +34031,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -37246,7 +37277,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -40465,7 +40501,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -43700,7 +43741,12 @@ SIMDJSON_NO_SANITIZE_UNDEFINED simdjson_inline void parse_integer_digits(const uint8_t *&p, uint64_t &i) { #ifdef SIMDJSON_SWAR_NUMBER_PARSING #if SIMDJSON_SWAR_NUMBER_PARSING - const uint8_t *const swar_end = p + 16; + // Identifiers, timestamps and counters often have eight digits or more. + if (is_made_of_eight_digits_fast(p)) { + i = i * 100000000 + parse_eight_digits_unrolled(p); + p += 8; + } + const uint8_t *const swar_end = p + 8; while (p < swar_end && is_made_of_four_digits_fast(p)) { i = i * 10000 + parse_four_digits_unrolled(p); p += 4; @@ -46666,8 +46712,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -46680,8 +46733,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -46689,7 +46748,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -46738,14 +46797,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -46779,8 +46864,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -46797,8 +46882,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -46816,19 +46901,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -46841,9 +46931,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -46865,12 +46955,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -46886,16 +46976,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -46918,8 +47010,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -46931,9 +47025,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -46948,7 +47042,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -46968,7 +47066,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -46978,7 +47076,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -46994,9 +47092,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -47007,9 +47105,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -47020,9 +47118,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -47044,12 +47142,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -47071,6 +47169,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -47096,17 +47338,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -47115,17 +47353,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -47140,18 +47374,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -47162,22 +47392,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -47237,25 +47483,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace arm64 // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = arm64::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - arm64::builder::string_builder b(initial_capacity); - arm64::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return arm64::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = arm64::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - arm64::builder::string_builder b(initial_capacity); - arm64::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return arm64::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -47460,6 +47699,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -47968,12 +48210,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -48145,96 +48413,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -49231,8 +49538,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -49245,8 +49559,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -49254,7 +49574,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -49303,14 +49623,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -49344,8 +49690,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -49362,8 +49708,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -49381,19 +49727,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -49406,9 +49757,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -49430,12 +49781,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -49451,16 +49802,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -49483,8 +49836,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -49496,9 +49851,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -49513,7 +49868,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -49533,7 +49892,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -49543,7 +49902,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -49559,9 +49918,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -49572,9 +49931,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -49585,9 +49944,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -49609,12 +49968,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -49636,6 +49995,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -49661,17 +50164,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -49680,17 +50179,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -49705,18 +50200,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -49727,22 +50218,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -49802,25 +50309,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace fallback // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = fallback::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - fallback::builder::string_builder b(initial_capacity); - fallback::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return fallback::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = fallback::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - fallback::builder::string_builder b(initial_capacity); - fallback::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return fallback::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -50025,6 +50525,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -50533,12 +51036,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -50710,96 +51239,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -52273,8 +52841,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -52287,8 +52862,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -52296,7 +52877,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -52345,14 +52926,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -52386,8 +52993,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -52404,8 +53011,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -52423,19 +53030,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -52448,9 +53060,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -52472,12 +53084,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -52493,16 +53105,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -52525,8 +53139,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -52538,9 +53154,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -52555,7 +53171,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -52575,7 +53195,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -52585,7 +53205,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -52601,9 +53221,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -52614,9 +53234,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -52627,9 +53247,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -52651,12 +53271,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -52678,6 +53298,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -52703,17 +53467,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -52722,17 +53482,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -52747,18 +53503,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -52769,22 +53521,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -52844,25 +53612,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace haswell // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = haswell::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - haswell::builder::string_builder b(initial_capacity); - haswell::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return haswell::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = haswell::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - haswell::builder::string_builder b(initial_capacity); - haswell::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return haswell::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -53067,6 +53828,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -53575,12 +54339,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -53752,96 +54542,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -55315,8 +56144,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -55329,8 +56165,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -55338,7 +56180,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -55387,14 +56229,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -55428,8 +56296,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -55446,8 +56314,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -55465,19 +56333,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -55490,9 +56363,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -55514,12 +56387,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -55535,16 +56408,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -55567,8 +56442,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -55580,9 +56457,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -55597,7 +56474,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -55617,7 +56498,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -55627,7 +56508,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -55643,9 +56524,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -55656,9 +56537,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -55669,9 +56550,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -55693,12 +56574,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -55720,6 +56601,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -55745,17 +56770,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -55764,17 +56785,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -55789,18 +56806,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -55811,22 +56824,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -55886,25 +56915,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace icelake // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = icelake::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - icelake::builder::string_builder b(initial_capacity); - icelake::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return icelake::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = icelake::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - icelake::builder::string_builder b(initial_capacity); - icelake::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return icelake::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -56109,6 +57131,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -56617,12 +57642,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -56794,96 +57845,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -58472,8 +59562,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -58486,8 +59583,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -58495,7 +59598,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -58544,14 +59647,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -58585,8 +59714,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -58603,8 +59732,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -58622,19 +59751,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -58647,9 +59781,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -58671,12 +59805,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -58692,16 +59826,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -58724,8 +59860,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -58737,9 +59875,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -58754,7 +59892,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -58774,7 +59916,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -58784,7 +59926,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -58800,9 +59942,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -58813,9 +59955,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -58826,9 +59968,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -58850,12 +59992,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -58877,6 +60019,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -58902,17 +60188,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -58921,17 +60203,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -58946,18 +60224,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -58968,22 +60242,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -59043,25 +60333,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace ppc64 // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = ppc64::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - ppc64::builder::string_builder b(initial_capacity); - ppc64::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return ppc64::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = ppc64::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - ppc64::builder::string_builder b(initial_capacity); - ppc64::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return ppc64::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -59266,6 +60549,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -59774,12 +61060,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -59951,96 +61263,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -61936,8 +63287,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -61950,8 +63308,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -61959,7 +63323,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -62008,14 +63372,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -62049,8 +63439,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -62067,8 +63457,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -62086,19 +63476,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -62111,9 +63506,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -62135,12 +63530,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -62156,16 +63551,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -62188,8 +63585,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -62201,9 +63600,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -62218,7 +63617,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -62238,7 +63641,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -62248,7 +63651,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -62264,9 +63667,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -62277,9 +63680,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -62290,9 +63693,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -62314,12 +63717,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -62341,6 +63744,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -62366,17 +63913,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -62385,17 +63928,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -62410,18 +63949,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -62432,22 +63967,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -62507,25 +64058,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace westmere // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = westmere::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - westmere::builder::string_builder b(initial_capacity); - westmere::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return westmere::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = westmere::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - westmere::builder::string_builder b(initial_capacity); - westmere::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return westmere::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -62730,6 +64274,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -63238,12 +64785,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -63415,96 +64988,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -64890,8 +66502,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -64904,8 +66523,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -64913,7 +66538,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -64962,14 +66587,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -65003,8 +66654,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -65021,8 +66672,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -65040,19 +66691,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -65065,9 +66721,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -65089,12 +66745,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -65110,16 +66766,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -65142,8 +66800,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -65155,9 +66815,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -65172,7 +66832,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -65192,7 +66856,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -65202,7 +66866,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -65218,9 +66882,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -65231,9 +66895,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -65244,9 +66908,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -65268,12 +66932,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -65295,6 +66959,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -65320,17 +67128,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -65339,17 +67143,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -65364,18 +67164,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -65386,22 +67182,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -65461,25 +67273,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace lsx // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = lsx::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - lsx::builder::string_builder b(initial_capacity); - lsx::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return lsx::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = lsx::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - lsx::builder::string_builder b(initial_capacity); - lsx::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return lsx::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -65684,6 +67489,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -66192,12 +68000,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -66369,96 +68203,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -67867,8 +69740,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -67881,8 +69761,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -67890,7 +69776,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -67939,14 +69825,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -67980,8 +69892,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -67998,8 +69910,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -68017,19 +69929,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -68042,9 +69959,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -68066,12 +69983,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -68087,16 +70004,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -68119,8 +70038,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -68132,9 +70053,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -68149,7 +70070,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -68169,7 +70094,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -68179,7 +70104,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -68195,9 +70120,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -68208,9 +70133,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -68221,9 +70146,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -68245,12 +70170,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -68272,6 +70197,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -68297,17 +70366,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -68316,17 +70381,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -68341,18 +70402,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -68363,22 +70420,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -68438,25 +70511,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace lasx // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = lasx::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - lasx::builder::string_builder b(initial_capacity); - lasx::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return lasx::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = lasx::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - lasx::builder::string_builder b(initial_capacity); - lasx::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return lasx::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -68661,6 +70727,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -69169,12 +71238,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -69346,96 +71441,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -70847,8 +72981,15 @@ namespace builder { // name lookup falls back to the wrong outer namespace). namespace internal { simdjson_really_inline char *write_uint_jeaiii(char *p, uint64_t v) noexcept; +simdjson_inline char *write_double(char *p, double v) noexcept; } // namespace internal -inline size_t write_string_escaped(const std::string_view input, char *out); +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out); + +SIMDJSON_PUSH_DISABLE_WARNINGS +SIMDJSON_DISABLE_GCC_WARNING(-Warray-bounds) +#if !defined(__clang__) +SIMDJSON_DISABLE_GCC_WARNING(-Wstringop-overflow) +#endif // ============================================================= // `writer`: position-as-local hot-path writer used by the reflection @@ -70861,8 +73002,14 @@ inline size_t write_string_escaped(const std::string_view input, char *out); // breaks the strict-aliasing penalty on every char* write through the // buffer, which forces a reload of `b.position` and `b.capacity` // after every byte. +// +// basic_writer (below) is the unchecked variant: the caller has +// already reserved enough capacity for everything the write chain can +// produce (see bound_detail::size_bound), so ensure() compiles away. // ============================================================= -struct writer { +template +struct basic_writer { + static constexpr bool checked = Checked; char *ptr; // buffer pointer (refreshed after a grow) size_t pos; // write position (local) size_t cap; // capacity (refreshed after a grow) @@ -70870,7 +73017,7 @@ struct writer { // Snapshot string_builder state into a writer for the duration of // a write chain. - simdjson_really_inline writer(string_builder &builder) noexcept + simdjson_really_inline basic_writer(string_builder &builder) noexcept : ptr(builder.unsafe_data()) , pos(builder.unsafe_position()) , cap(builder.unsafe_capacity()) @@ -70919,14 +73066,40 @@ struct writer { } }; +// The unchecked writer writes into a raw buffer that the caller sized with +// serialized_size_bound: it never grows and needs no string_builder. +template <> +struct basic_writer { + static constexpr bool checked = false; + char *ptr; + size_t pos; + + simdjson_really_inline basic_writer(char *buffer, size_t position) noexcept + : ptr(buffer), pos(position) {} + + simdjson_really_inline bool ensure(size_t) const noexcept { return true; } +}; + +using writer = basic_writer; +using unchecked_writer = basic_writer; + +// Bytes reserved past the size bound for an unchecked writer: it may then +// write a little past the end of what it produces (e.g., copy keys as whole +// 16-byte blocks). +inline constexpr size_t unchecked_slack = 64; + +consteval size_t padded_key_length(size_t length) { + return (length + 15) / 16 * 16; +} + // === Helper: invoke a string_builder member that writes variable-length // content (escape_and_append_with_quotes etc), syncing the writer's local // state before the call and reloading after. Used for string fields where // rewriting the entire SIMD escape path through the writer would be a much // bigger refactor. f may be user code (a with serializer) that // throws: the exception then propagates to the caller. -template -simdjson_really_inline void call_through_string_builder(writer &w, F &&f) noexcept(noexcept(f(w.sb))) { +template +simdjson_really_inline void call_through_string_builder(W &w, F &&f) noexcept(noexcept(f(w.sb))) { w.sync(); f(w.sb); w.ptr = w.sb.unsafe_data(); @@ -70960,8 +73133,8 @@ simdjson_really_inline bool should_serialize(const V &value) { // Serialize a member value, through its with annotation when the // adapter provides a serialize function. -template -simdjson_really_inline void atom_member(writer &w, const V &value) { +template +simdjson_really_inline void atom_member(W &w, const V &value) { constexpr std::meta::info with_type = simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t); if constexpr (with_type != std::meta::info{}) { using adapter = typename [: with_type :]::adapter; @@ -70978,8 +73151,8 @@ simdjson_really_inline void atom_member(writer &w, const V &value) { // Write the "key":value pairs of the members of t (without the braces), each // preceded by a comma unless it is the first one. The members of a member // annotated with flatten are written in its place. -template -simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { +template +simdjson_really_inline void atom_fields(W &w, const T &t, bool &first) { // Per-field block: ensure key+value worst case, then write key + value // through the writer's local pos. For arithmetic fields, the integer // write happens directly via write_uint_jeaiii on w.ptr+w.pos, so pos @@ -70997,19 +73170,24 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { "simdjson::flatten requires a member whose type is a structure serialized member by member"); atom_fields(w, t.[:dm:], first); } else { + // Copy the key as whole 16-byte blocks from a zero-padded copy (one + // load and one store); ensure() reserves the padded length, and the + // unchecked writer has slack past its bound. constexpr const char* key_name = simdjson::get_json_key_name(); + constexpr size_t first_key_len = constevalutil::consteval_to_quoted_escaped(key_name).size() + 1; + constexpr size_t rest_key_len = first_key_len + 1; constexpr auto first_key = std::define_static_string( - constevalutil::consteval_to_quoted_escaped(key_name) + ":"); + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(first_key_len) - first_key_len, '\0')); constexpr auto rest_key = std::define_static_string( - std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":"); - constexpr size_t first_key_len = std::char_traits::length(first_key); - constexpr size_t rest_key_len = std::char_traits::length(rest_key); - if (!w.ensure(rest_key_len)) { return; } + std::string(",") + constevalutil::consteval_to_quoted_escaped(key_name) + ":" + + std::string(padded_key_length(rest_key_len) - rest_key_len, '\0')); + if (!w.ensure(padded_key_length(rest_key_len))) { return; } if (first) { - std::memcpy(w.ptr + w.pos, first_key, first_key_len); + std::memcpy(w.ptr + w.pos, first_key, padded_key_length(first_key_len)); w.pos += first_key_len; } else { - std::memcpy(w.ptr + w.pos, rest_key, rest_key_len); + std::memcpy(w.ptr + w.pos, rest_key, padded_key_length(rest_key_len)); w.pos += rest_key_len; } first = false; @@ -71022,9 +73200,9 @@ simdjson_really_inline void atom_fields(writer &w, const T &t, bool &first) { } // namespace annotation_detail -template +template requires(concepts::container_but_not_string && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { auto it = t.begin(); auto end = t.end(); if (it == end) { @@ -71046,12 +73224,12 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { w.ptr[w.pos++] = ']'; } -template +template requires(std::is_same_v || std::is_same_v || std::is_same_v || std::is_same_v) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { // Inline the escape path through the writer so we never round-trip // pos through memory for string fields (Twitter is dominated by // these -- sync/reload around each string was a real cost). @@ -71067,16 +73245,18 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * input.size())) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(input.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * input.size())) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(input, w.ptr + w.pos); w.ptr[w.pos++] = '"'; } -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &m) { +simdjson_really_inline constexpr void atom(W &w, const T &m) { if (m.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "{}", 2); @@ -71099,8 +73279,10 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { // subsequent escape would overflow the buffer. max - w.pos cannot wrap, and // size < (max - pos) / 6 implies pos + 6 * size + 6 <= max. // Note that this is pedantic except maybe on 32-bit targets. - if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } - if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + if constexpr (W::checked) { + if (simdjson_unlikely(key_sv.size() >= ((std::numeric_limits::max)() - w.pos) / 6)) { return; } + if (!w.ensure(2 + 6 * key_sv.size() + 1)) { return; } + } w.ptr[w.pos++] = '"'; w.pos += write_string_escaped(key_sv, w.ptr + w.pos); w.ptr[w.pos++] = '"'; @@ -71112,9 +73294,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &m) { } -template::value && !std::is_same_v>::type> -simdjson_really_inline constexpr void atom(writer &w, const number_type t) { +simdjson_really_inline constexpr void atom(W &w, const number_type t) { // Booleans / floats: defer to string_builder (rare path; keeps writer hot // path free of float-formatter machinery). For integers, write directly // via jeaiii using local pos. @@ -71129,7 +73311,11 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { w.pos += 5; } } else if constexpr (std::is_floating_point_v) { - call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + if constexpr (W::checked) { + call_through_string_builder(w, [&](string_builder &b) { b.append(t); }); + } else { + w.pos = size_t(internal::write_double(w.ptr + w.pos, double(t)) - w.ptr); + } } else if constexpr (std::is_unsigned_v) { if (!w.ensure(20)) return; char *end = internal::write_uint_jeaiii( @@ -71149,7 +73335,7 @@ simdjson_really_inline constexpr void atom(writer &w, const number_type t) { } } -template +template requires(std::is_class_v && !concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && @@ -71159,7 +73345,7 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &t) { +simdjson_really_inline constexpr void atom(W &w, const T &t) { if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { // A transparent structure is serialized as its single member. constexpr auto dm = simdjson::detail::transparent_member(^^T); @@ -71175,9 +73361,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &t) { } // Support for optional types (std::optional, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &opt) { +simdjson_really_inline constexpr void atom(W &w, const T &opt) { if (opt) { atom(w, opt.value()); } else { @@ -71188,9 +73374,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &opt) { } // Support for smart pointers (std::unique_ptr, std::shared_ptr, etc.) -template +template requires(!require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { +simdjson_really_inline constexpr void atom(W &w, const T &ptr) { if (ptr) { atom(w, *ptr); } else { @@ -71201,9 +73387,9 @@ simdjson_really_inline constexpr void atom(writer &w, const T &ptr) { } // Support for enums - serialize as string representation using expand approach from P2996R12 -template +template requires(std::is_enum_v && !require_custom_serialization) -simdjson_really_inline void atom(writer &w, const T &e) { +simdjson_really_inline void atom(W &w, const T &e) { #if SIMDJSON_STATIC_REFLECTION static constexpr auto enumerators = std::define_static_array(std::meta::enumerators_of(^^T)); template for (constexpr auto enum_val : enumerators) { @@ -71225,12 +73411,12 @@ simdjson_really_inline void atom(writer &w, const T &e) { } // Support for appendable containers that don't have operator[] (sets, etc.) -template +template requires(!concepts::container_but_not_string && !concepts::string_view_keyed_map && !concepts::optional_type && !concepts::smart_pointer && !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) -simdjson_really_inline constexpr void atom(writer &w, const T &container) { +simdjson_really_inline constexpr void atom(W &w, const T &container) { if (container.empty()) { if (!w.ensure(2)) return; std::memcpy(w.ptr + w.pos, "[]", 2); @@ -71252,6 +73438,150 @@ simdjson_really_inline constexpr void atom(writer &w, const T &container) { w.ptr[w.pos++] = ']'; } +// ============================================================= +// Size bound: an upper bound on the number of bytes that atom(w, t) writes. +// Computing it first lets append() reserve the capacity once and then run +// the whole write chain through an unchecked_writer, without a capacity +// check before every write. It mirrors the atom() overloads above. +// ============================================================= +namespace bound_detail { + +// Whether size_bound covers everything that atom() writes for T: not when a +// member is serialized by a with serializer, which writes an unknown +// amount through the string_builder. +template +consteval bool is_bounded() { + if constexpr (require_custom_serialization) { + return false; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v || std::is_arithmetic_v || std::is_enum_v) { + return true; + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return is_bounded())>>(); + } else if constexpr (concepts::string_view_keyed_map) { + return is_bounded>(); + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + return is_bounded>>(); + } else { + bool bounded = true; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + bounded = bounded && simdjson::detail::annotation_of_template(dm, ^^simdjson::detail::with_t) == std::meta::info{} && + is_bounded().[:dm:])>>(); + } + }; + return bounded; + } +} + +template +consteval size_t enum_bound() { + size_t bound = 20; // the integer fallback + template for (constexpr auto enum_val : std::define_static_array(std::meta::enumerators_of(^^T))) { + constexpr size_t len = std::char_traits::length(std::define_static_string( + constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()))); + bound = (std::max)(bound, len); + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound(const T &t) noexcept; + +// Bound for the "key":value pairs of a structure, commas included. +template +simdjson_really_inline size_t fields_bound(const T &t) noexcept { + size_t bound = 0; + template for (constexpr auto dm : std::define_static_array(std::meta::nonstatic_data_members_of(^^T, std::meta::access_context::unchecked()))) { + if constexpr (annotation_detail::is_serialized_member(dm)) { + if constexpr (simdjson::detail::has_annotation(dm, ^^simdjson::detail::flatten_tag)) { + bound += fields_bound(t.[:dm:]); + } else { + constexpr size_t rest_key_len = constevalutil::consteval_to_quoted_escaped(simdjson::get_json_key_name()).size() + 2; + bound += rest_key_len + size_bound(t.[:dm:]); + } + } + }; + return bound; +} + +template +simdjson_really_inline size_t size_bound([[maybe_unused]] const T &t) noexcept { + if constexpr (std::is_same_v) { + return 2 + 6; + } else if constexpr (std::is_same_v || std::is_same_v || + std::is_same_v) { + // Every byte may become \uXXXX, plus the quotes. + return 2 + 6 * std::string_view(t).size(); + } else if constexpr (std::is_same_v) { + return 5; + } else if constexpr (std::is_floating_point_v) { + return simdjson::internal::to_chars_buffer_size; + } else if constexpr (std::is_arithmetic_v) { + return 20; + } else if constexpr (std::is_enum_v) { + return enum_bound(); + } else if constexpr (concepts::optional_type || concepts::smart_pointer) { + return t ? size_bound(*t) : 4; + } else if constexpr (concepts::string_view_keyed_map) { + size_t bound = 2; + for (const auto &[key, value] : t) { + // comma, quotes, colon + bound += 4 + 6 * std::string_view(key).size() + size_bound(value); + } + return bound; + } else if constexpr (concepts::container_but_not_string || concepts::appendable_containers) { + using value_type = std::remove_cvref_t>; + if constexpr (std::is_arithmetic_v && !std::is_same_v) { + // A fixed bound per element: no need to visit them. + return 2 + size_t(std::ranges::distance(t)) * (1 + size_bound(value_type{})); + } else { + size_t bound = 2; +#if defined(__GNUC__) && !defined(__clang__) +#pragma GCC novector // a vector loop is slower on short containers +#endif +#pragma GCC unroll 4 + for (const auto &item : t) { + bound += 1 + size_bound(item); + } + return bound; + } + } else if constexpr (simdjson::detail::has_annotation(^^T, ^^simdjson::detail::transparent_tag)) { + constexpr auto dm = simdjson::detail::transparent_member(^^T); + return size_bound(t.[:dm:]); + } else { + return 2 + fields_bound(t); + } +} + +} // namespace bound_detail + +// Write t through an unchecked writer when its size bound is available, +// reserving that many bytes first, and through the checked writer otherwise. +template +simdjson_really_inline void append_bounded(string_builder &b, const T &t) { + // On 32-bit systems, the bound could overflow: keep the checked writer. + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + const size_t bound = bound_detail::size_bound(t) + unchecked_slack; + const size_t pos = b.unsafe_position(); + // The bound is a sum of in-memory sizes times a small constant: it cannot + // overflow on a 64-bit system. Be pedantic elsewhere. + if (sizeof(size_t) >= 8 || bound <= (std::numeric_limits::max)() - pos) { + const size_t cap = b.unsafe_capacity(); + // Grow geometrically so that many small appends stay amortized. + if (pos + bound <= cap || b.unsafe_grow((std::max)(cap * 2, pos + bound))) { + unchecked_writer w(b.unsafe_data(), pos); + atom(w, t); + b.unsafe_set_position(w.pos); + } + return; + } + } + writer w(b); + atom(w, t); + w.sync(); +} + // append() -- top-level entry. Each overload constructs a stack-local // writer, runs atom(w, t) through the inlined call chain, then syncs // the local position back into the string_builder. @@ -71277,17 +73607,13 @@ simdjson_inline void append(string_builder &b, const T &t) { template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template @@ -71296,17 +73622,13 @@ template !std::is_same_v && !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } template requires(!require_custom_serialization) simdjson_inline void append(string_builder &b, const T &t) { - writer w(b); - atom(w, t); - w.sync(); + append_bounded(b, t); } // works for struct @@ -71321,18 +73643,14 @@ template !std::is_same_v && !std::is_same_v && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } // works for container that have begin() and end() iterators template requires(concepts::container_but_not_string && !require_custom_serialization) simdjson_inline void append(string_builder &b, const Z &z) { - writer w(b); - atom(w, z); - w.sync(); + append_bounded(b, z); } template @@ -71343,22 +73661,38 @@ void append(string_builder &b, const Z &z) { template -simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); +simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + if constexpr (sizeof(size_t) >= 8 && bound_detail::is_bounded()) { + // Write straight into s, sized by the bound: no intermediate buffer, no copy. + (void)initial_capacity; + const size_t bound = bound_detail::size_bound(z) + unchecked_slack; + auto write = [&z](char *p) noexcept { + unchecked_writer w(p, 0); + atom(w, z); + return w.pos; + }; +#if defined(__cpp_lib_string_resize_and_overwrite) && __cpp_lib_string_resize_and_overwrite >= 202110L + s.resize_and_overwrite(bound, [&write](char *p, size_t) noexcept { return write(p); }); +#else + s.resize(bound); + s.resize(write(s.data())); +#endif + return SUCCESS; + } else { + string_builder b(initial_capacity); + append(b, z); + std::string_view view; + if(auto e = b.view().get(view); e) { return e; } + s.assign(view); + return SUCCESS; + } } template -simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { - string_builder b(initial_capacity); - append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; +simdjson_warn_unused simdjson_result to_json_string(const Z &z, size_t initial_capacity = string_builder::DEFAULT_INITIAL_CAPACITY) { + std::string s; + if(auto e = to_json(z, s, initial_capacity); e) { return e; } + return s; } template @@ -71418,25 +73752,18 @@ simdjson_warn_unused simdjson_result extract_from(const T &obj, siz return std::string(s); } +SIMDJSON_POP_DISABLE_WARNINGS + } // namespace builder } // namespace rvv_vls // Alias the function template to 'to' in the global namespace template simdjson_warn_unused simdjson_result to_json(const Z &z, size_t initial_capacity = rvv_vls::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - rvv_vls::builder::string_builder b(initial_capacity); - rvv_vls::builder::append(b, z); - std::string_view s; - if(auto e = b.view().get(s); e) { return e; } - return std::string(s); + return rvv_vls::builder::to_json_string(z, initial_capacity); } template simdjson_warn_unused error_code to_json(const Z &z, std::string &s, size_t initial_capacity = rvv_vls::builder::string_builder::DEFAULT_INITIAL_CAPACITY) { - rvv_vls::builder::string_builder b(initial_capacity); - rvv_vls::builder::append(b, z); - std::string_view view; - if(auto e = b.view().get(view); e) { return e; } - s.assign(view); - return SUCCESS; + return rvv_vls::builder::to_json(z, s, initial_capacity); } // Global namespace function for extract_from template @@ -71641,6 +73968,9 @@ simdjson_warn_unused simdjson_result extract_fractured_json( #endif #if SIMDJSON_EXPERIMENTAL_HAS_SSE2 #include +#if defined(__AVX2__) +#include +#endif #ifdef _MSC_VER #include #endif @@ -72149,12 +74479,38 @@ simdjson_never_inline char *escape_block(const uint8_t *src, char *out, // Writes the escaped version of input to out, returning the number of bytes // written. -inline size_t write_string_escaped(const std::string_view input, char *out) { +simdjson_really_inline size_t write_string_escaped(const std::string_view input, char *out) { const size_t len = input.size(); const uint8_t *src = reinterpret_cast(input.data()); const char *const initout = out; size_t i = 0; +#if SIMDJSON_EXPERIMENTAL_HAS_SSE2 && defined(__AVX2__) + while (i + 32 <= len) { + const __m256i word = _mm256_loadu_si256(reinterpret_cast(src + i)); + const __m256i flags = _mm256_or_si256( + _mm256_or_si256(_mm256_cmpeq_epi8(word, _mm256_set1_epi8(34)), // '"' + _mm256_cmpeq_epi8(word, _mm256_set1_epi8(92))), // '\\' + _mm256_cmpeq_epi8(_mm256_subs_epu8(word, _mm256_set1_epi8(31)), + _mm256_setzero_si256())); // control + const uint32_t mask = uint32_t(_mm256_movemask_epi8(flags)); + if (simdjson_likely(mask == 0)) { + _mm256_storeu_si256(reinterpret_cast<__m256i *>(out), word); + out += 32; + } else { + for (size_t half = 0; half < 32; half += 16) { + const uint64_t m = (mask >> half) & 0xFFFF; + if (m == 0) { + escape_store16(out, escape_load16(src + i + half)); + out += 16; + } else { + out = escape_block(src, out, i + half, i + half + 16, m); + } + } + } + i += 32; + } +#endif while (i + 16 <= len) { escape_vector word = escape_load16(src + i); escape_vector flags = escape_flags(word); @@ -72326,96 +74682,135 @@ simdjson_inline void string_builder::clear() noexcept { namespace internal { -static const char decimal_table[200] = { - 0x30, 0x30, 0x30, 0x31, 0x30, 0x32, 0x30, 0x33, 0x30, 0x34, 0x30, 0x35, - 0x30, 0x36, 0x30, 0x37, 0x30, 0x38, 0x30, 0x39, 0x31, 0x30, 0x31, 0x31, - 0x31, 0x32, 0x31, 0x33, 0x31, 0x34, 0x31, 0x35, 0x31, 0x36, 0x31, 0x37, - 0x31, 0x38, 0x31, 0x39, 0x32, 0x30, 0x32, 0x31, 0x32, 0x32, 0x32, 0x33, - 0x32, 0x34, 0x32, 0x35, 0x32, 0x36, 0x32, 0x37, 0x32, 0x38, 0x32, 0x39, - 0x33, 0x30, 0x33, 0x31, 0x33, 0x32, 0x33, 0x33, 0x33, 0x34, 0x33, 0x35, - 0x33, 0x36, 0x33, 0x37, 0x33, 0x38, 0x33, 0x39, 0x34, 0x30, 0x34, 0x31, - 0x34, 0x32, 0x34, 0x33, 0x34, 0x34, 0x34, 0x35, 0x34, 0x36, 0x34, 0x37, - 0x34, 0x38, 0x34, 0x39, 0x35, 0x30, 0x35, 0x31, 0x35, 0x32, 0x35, 0x33, - 0x35, 0x34, 0x35, 0x35, 0x35, 0x36, 0x35, 0x37, 0x35, 0x38, 0x35, 0x39, - 0x36, 0x30, 0x36, 0x31, 0x36, 0x32, 0x36, 0x33, 0x36, 0x34, 0x36, 0x35, - 0x36, 0x36, 0x36, 0x37, 0x36, 0x38, 0x36, 0x39, 0x37, 0x30, 0x37, 0x31, - 0x37, 0x32, 0x37, 0x33, 0x37, 0x34, 0x37, 0x35, 0x37, 0x36, 0x37, 0x37, - 0x37, 0x38, 0x37, 0x39, 0x38, 0x30, 0x38, 0x31, 0x38, 0x32, 0x38, 0x33, - 0x38, 0x34, 0x38, 0x35, 0x38, 0x36, 0x38, 0x37, 0x38, 0x38, 0x38, 0x39, - 0x39, 0x30, 0x39, 0x31, 0x39, 0x32, 0x39, 0x33, 0x39, 0x34, 0x39, 0x35, - 0x39, 0x36, 0x39, 0x37, 0x39, 0x38, 0x39, 0x39, -}; - -// Forward unsigned-int writer (cascade-on-magnitude, no upfront digit_count). -// Built from a non-recursive DAG of always_inline helpers -- gcc and MSVC -// refuse to inline recursive `always_inline`/`__forceinline` functions. -// Caller must guarantee at least 20 bytes available at p. All helpers -// return pointer past the last digit written. - -// Caller guarantees v < 100. Writes 1-2 digits. -simdjson_really_inline char* write_lt100(char* p, uint64_t v) noexcept { - if (v < 10) { *p++ = char('0' + v); return p; } - std::memcpy(p, &decimal_table[v * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Writes 1-4 digits. -simdjson_really_inline char* write_lt10000(char* p, uint64_t v) noexcept { - if (v < 100) return write_lt100(p, v); - uint64_t hi = v / 100, lo = v % 100; - if (v < 1000) { - *p++ = char('0' + hi); - } else { - std::memcpy(p, &decimal_table[hi * 2], 2); - p += 2; - } - std::memcpy(p, &decimal_table[lo * 2], 2); - return p + 2; -} - -// Caller guarantees v < 10000. Always writes exactly 4 digits. -simdjson_really_inline void write_4_digits(char* p, uint64_t v) noexcept { - uint64_t hi = v / 100, lo = v % 100; - std::memcpy(p, &decimal_table[hi * 2], 2); - std::memcpy(p + 2, &decimal_table[lo * 2], 2); -} - -// Caller guarantees v < 10^8. Writes 1-8 digits. -simdjson_really_inline char* write_lt1e8(char* p, uint64_t v) noexcept { - if (v < 10000) return write_lt10000(p, v); - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; -} - -simdjson_really_inline char* write_uint_jeaiii(char* p, uint64_t v) noexcept { - if (v < 10000ULL) return write_lt10000(p, v); - if (v < 100000000ULL) { // 5-8 digits - uint64_t hi = v / 10000, lo = v % 10000; - p = write_lt10000(p, hi); - write_4_digits(p, lo); - return p + 4; - } - if (v < 10000000000000000ULL) { // 9-16 digits - uint64_t hi = v / 100000000ULL, lo = v % 100000000ULL; - p = write_lt1e8(p, hi); - uint64_t lo_hi = lo / 10000, lo_lo = lo % 10000; - write_4_digits(p, lo_hi); - write_4_digits(p + 4, lo_lo); +// Integer to decimal: James Edward Anhalt III's algorithm +static const char jeaiii_dd[201] = + "00010203040506070809101112131415161718192021222324252627282930313233343536373839" + "40414243444546474849505152535455565758596061626364656667686970717273747576777879" + "8081828384858687888990919293949596979899"; +static const char jeaiii_fd[201] = + "0\0" "1\0" "2\0" "3\0" "4\0" "5\0" "6\0" "7\0" "8\0" "9\0" + "10111213141516171819202122232425262728293031323334353637383940414243444546474849" + "50515253545556575859606162636465666768697071727374757677787980818283848586878889" + "90919293949596979899"; + +simdjson_really_inline void jeaiii_write_dd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_dd[2 * k], 2); +} +simdjson_really_inline void jeaiii_write_fd(char *p, uint64_t k) noexcept { + std::memcpy(p, &jeaiii_fd[2 * k], 2); +} + +// Caller guarantees n < 10^8. Writes 1 to 8 digits. +simdjson_really_inline char *jeaiii_lt1e8(char *b, uint32_t n) noexcept { + constexpr uint64_t mask24 = (uint64_t(1) << 24) - 1; + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + if (n < 100) { + jeaiii_write_fd(b, n); + return n < 10 ? b + 1 : b + 2; + } + if (n < 1000000) { + if (n < 10000) { + const uint32_t f0 = uint32_t(10 * (1 << 24) / 1e3 + 1) * n; + jeaiii_write_fd(b, f0 >> 24); + b -= n < 1000; + const uint32_t f2 = uint32_t(f0 & mask24) * 100; + jeaiii_write_dd(b + 2, f2 >> 24); + return b + 4; + } + const uint64_t f0 = uint64_t(10 * (1ull << 32) / 1e5 + 1) * n; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 100000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + return b + 6; + } + const uint64_t f0 = uint64_t(10 * (1ull << 48) / 1e7 + 1) * n >> 16; + jeaiii_write_fd(b, f0 >> 32); + b -= n < 10000000; + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees z < 10^8. Always writes exactly 8 digits. +simdjson_really_inline char *jeaiii_8_digits(char *b, uint32_t z) noexcept { + constexpr uint64_t mask32 = (uint64_t(1) << 32) - 1; + const uint64_t f0 = (uint64_t((1ull << 48) / 1e6 + 1) * z >> 16) + 1; + jeaiii_write_dd(b, f0 >> 32); + const uint64_t f2 = (f0 & mask32) * 100; + jeaiii_write_dd(b + 2, f2 >> 32); + const uint64_t f4 = (f2 & mask32) * 100; + jeaiii_write_dd(b + 4, f4 >> 32); + const uint64_t f6 = (f4 & mask32) * 100; + jeaiii_write_dd(b + 6, f6 >> 32); + return b + 8; +} + +// Caller guarantees 10^8 <= n < 2^32. Writes 9 or 10 digits. +simdjson_really_inline char *jeaiii_9_or_10(char *b, uint64_t n) noexcept { + constexpr uint64_t mask57 = (uint64_t(1) << 57) - 1; + const uint64_t f0 = uint64_t(10 * (1ull << 57) / 1e9 + 1) * n; + jeaiii_write_fd(b, f0 >> 57); + b -= n < 1000000000; + const uint64_t f2 = (f0 & mask57) * 100; + jeaiii_write_dd(b + 2, f2 >> 57); + const uint64_t f4 = (f2 & mask57) * 100; + jeaiii_write_dd(b + 4, f4 >> 57); + const uint64_t f6 = (f4 & mask57) * 100; + jeaiii_write_dd(b + 6, f6 >> 57); + const uint64_t f8 = (f6 & mask57) * 100; + jeaiii_write_dd(b + 8, f8 >> 57); + return b + 10; +} + +simdjson_really_inline char *write_uint_jeaiii(char *b, uint64_t n) noexcept { + if (n < 100000000) { + return jeaiii_lt1e8(b, uint32_t(n)); + } + if (n < (uint64_t(1) << 32)) { + return jeaiii_9_or_10(b, n); + } + // At least 10 digits: the low 8 digits, and 2 to 12 digits above them. + const uint32_t z = uint32_t(n % 100000000); + uint64_t u = n / 100000000; + if (u < 100000000) { + // u has 2 to 8 digits (if u < 10, n would be below 2^32). + b = jeaiii_lt1e8(b, uint32_t(u)); + } else if (u < (uint64_t(1) << 32)) { + b = jeaiii_9_or_10(b, u); + } else { + // u has 11 or 12 digits: split off 8 more. + const uint32_t y = uint32_t(u % 100000000); + u /= 100000000; + b = jeaiii_lt1e8(b, uint32_t(u)); // 3 or 4 digits + b = jeaiii_8_digits(b, y); + } + return jeaiii_8_digits(b, z); +} + +// Writes v at p, which must have to_chars_buffer_size bytes available, and +// returns the end of what was written. +simdjson_inline char *write_double(char *p, double v) noexcept { +#if SIMDJSON_ENABLE_NAN_INF + if (simdjson_unlikely(!std::isfinite(v))) { + if (std::isnan(v)) { + std::memcpy(p, "NaN", 3); + return p + 3; + } + if (v < 0) { + *p++ = '-'; + } + std::memcpy(p, "Infinity", 8); return p + 8; } - // 17-20 digits - uint64_t hi = v / 10000000000000000ULL, lo = v % 10000000000000000ULL; - p = write_lt10000(p, hi); - uint64_t lo_a = lo / 100000000ULL, lo_b = lo % 100000000ULL; - uint64_t lo_a_hi = lo_a / 10000, lo_a_lo = lo_a % 10000; - uint64_t lo_b_hi = lo_b / 10000, lo_b_lo = lo_b % 10000; - write_4_digits(p, lo_a_hi); - write_4_digits(p + 4, lo_a_lo); - write_4_digits(p + 8, lo_b_hi); - write_4_digits(p + 12, lo_b_lo); - return p + 16; +#endif + return simdjson::internal::to_chars(p, nullptr, v); } } // namespace internal @@ -81953,6 +84348,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -99616,6 +102052,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -117756,6 +120233,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -135896,6 +138414,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -154151,6 +156710,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -172713,6 +175313,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -190765,6 +193406,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -208840,6 +211522,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out); @@ -226918,6 +229641,47 @@ error_code tag_invoke(deserialize_tag, ValT &val, T &out) noexcept(false) { SIMDJSON_TRY(val.get_array().get(arr)); } + if constexpr (std::is_same_v> && !std::is_same_v) { + // Collect the elements in a per-thread scratch vector that keeps its + // capacity from call to call, then move them into out after reserving the + // exact size: out is allocated once instead of being regrown. A nested + // array of the same type finds the scratch busy and takes the paths below. + struct scratch_space { + std::vector elements{}; + bool busy{false}; + }; + static thread_local scratch_space scratch; + if (!scratch.busy && out.empty()) { + struct release_scratch { + scratch_space &s; + T &out; + size_t parsed{0}; + bool complete{false}; + // On an error or an exception, out gets the elements parsed so far (as + // with the loops below), without allocating. Kept out of the hot path. + simdjson_never_inline void keep_parsed() noexcept { + s.elements.resize(parsed); + out.swap(s.elements); + } + ~release_scratch() { + if (simdjson_unlikely(!complete)) { keep_parsed(); } + s.elements.clear(); + // Do not hold on to the memory of a very large array. + if (s.elements.capacity() * sizeof(value_type) > (1 << 20)) { std::vector().swap(s.elements); } + s.busy = false; + } + } release{scratch, out}; + scratch.busy = true; + for (auto v : arr) { + SIMDJSON_TRY(v.get(scratch.elements.emplace_back())); + release.parsed++; + } + out.reserve(release.parsed); + release.complete = true; + for (auto &e : scratch.elements) { out.emplace_back(std::move(e)); } + return SUCCESS; + } + } if constexpr (details::deserialize_in_place) { for (auto v : arr) { auto &slot = concepts::emplace_one(out);