DEBUG: field 482:11'html_meta_file': './data/etext/10556.html', @192151315
DEBUG: lookup expression for field 'html_meta_file'
DEBUG: got expression number for 'html_meta_file' to be 0
token positions of document '1055610556' are out or range (document too big, 150199 token positions assigned)
DEBUG: buffer reset, rest: 1055710557 Brooke, L. Leslie (Leonard Leslie), 1862-1940 Johnny Crow's Party English PZ: Language and Literatures:
failed to process document 'gutenberg.tsv': failed to process document 'gutenberg.tsv': error closing document in transaction: corrupt data (unpackInt32_ 1)
done
The positions of the experimental TSV segmenter with ZIP-file @zipinclude function are quite big,
because it's basically the position within the TSV file and the position of the file withing the
uncompressed ZIP stream.
The positions of the experimental TSV segmenter with ZIP-file @zipinclude function are quite big,
because it's basically the position within the TSV file and the position of the file withing the
uncompressed ZIP stream.
See
https://github.com/andreasbaumann/strusExamples/tree/master/gutenberg
and
https://github.com/andreasbaumann/strusAnalyzer/tree/tsv_extensions