Roadmap to bring petitparser-core to feature parity with the canonical Dart implementation while maintaining strict backward compatibility with existing code and adopting modern Java best practices (Java 21 baseline, standard functional interfaces, and stream idioms).
- Preserve all existing public methods, constructors, and classes.
- Downstream modules (
petitparser-json,petitparser-xml,petitparser-smalltalk) and user code must compile and pass all tests without modification. - Additive evolution: introduce new functionality alongside existing APIs (e.g.
starSeparatedreturningSeparatedListwithout removingseparatedByreturningList<Object>). - When introducing generic parameters (
Parser<R>,Result<R>, etc.), ensure raw-type usage compiles cleanly without warnings or breaking changes.
- Avoid mechanical translations of Dart-specific constructs (e.g. extension methods, positional record tuples, operator overloads).
- Leverage standard Java paradigms:
- Fluent method chaining on
Parserand builders. - Standard functional interfaces (
java.util.function.Function,Predicate,BiFunction,Supplier,Consumer). - Java
Streamand lazyIterable/Iteratorfor streaming matching and graph traversals. - Immutability for core state (
Context,Result,Token,SeparatedList) using defensive copies andCollections.unmodifiableList. - Null-safety checks using
Objects.requireNonNull(arg, "message").
- Fluent method chaining on
-
Zero-Allocation Fast Path: Every parser subclass must implement an optimized
int fastParseOn(String buffer, int position)that performs zero heap allocations (noContext,Result, or boxed wrapper instances) and returns-1on failure without throwing exceptions. -
Direct String Slicing (Zero-List Lexing): Character repeaters (
starString(),plusString(),repeatString()) must scan characters using primitive loop indexes onbuffer.charAt(i)and extract the final string with a single nativebuffer.substring(start, stop)call, completely bypassing intermediateList<Character>allocations. -
$O(1)$ Character Lookup Tables: Specialize character predicates with boolean lookup tables (boolean[256]) or primitive bitmasks (BitSet/long[]) for ASCII/Latin-1 character sets, executing in constant time rather than branching or binary searching. - Range Merging Optimization: Automatically sort and merge adjacent/overlapping character ranges during grammar construction so runtime checks execute the minimum possible comparisons.
-
HotSpot SIMD Intrinsics: String literal parsers must leverage JVM-intrinsic methods like
String.startsWith(prefix, position)andString.regionMatchesto take advantage of vectorized CPU instructions. -
Stateless Singleton Reuse: Use static singleton instances for stateless parsers and predicates (e.g.
PositionParser,EpsilonParser,NewlineParser,ConstantCharPredicate.any(),DigitCharPredicate.INSTANCE). -
Lazy Stream Matching: Matcher APIs must provide lazy
Stream<T>andIterable<T>pipelines to process matches on demand with minimal memory footprint.
-
Task 1.1: Marker & Lifecycle Interfaces
- Create
org.petitparser.parser.combinators.SequentialParsermarker interface. - Implement
SequentialParseronSequenceParserandTrimmingParser. - Create
org.petitparser.parser.combinators.ResolvableParserwithParser resolve(). - Implement
ResolvableParseronSettableParser.
- Create
-
Task 1.2: WhereParser & Filtering
- Implement
org.petitparser.parser.actions.WhereParser<T>withjava.util.function.Predicate<T>and custom error message / failure factoryBiFunction<Context, Result, Result>. - Add
where(Predicate<T>),where(Predicate<T>, String), andwhere(Predicate<T>, BiFunction<Context, Result, Result>)toParser. - Implement zero-allocation fast-parse, copy, and equality methods.
- Implement
-
Task 1.3: SkipParser & Delimiters
- Implement
org.petitparser.parser.combinators.SkipParserimplementingSequentialParser. - Add
skip(Parser before, Parser after)toParser. - Implement zero-allocation fast-parse, copy, child replacement, and equality methods.
- Implement
-
Task 1.4: LabelParser & Debugging
- Implement
org.petitparser.parser.combinators.LabelParserwrapping a delegate with a string label. - Add
labeled(String label)toParser. - Override
toString()to displaydelegate.toString() + "[" + label + "]".
- Implement
-
Task 1.5: PositionParser
- Implement
org.petitparser.parser.primitive.PositionParserreturningcontext.getPosition(). - Provide a reusable static singleton
PositionParser.INSTANCE. - Add static factory
Parser.position(). - Implement zero-allocation fast-parse returning position unchanged.
- Implement
-
Task 1.6: NewlineParser
- Implement
org.petitparser.parser.primitive.NewlineParsermatching\n,\r\n, and\r. - Add static factories
Parser.newline()andParser.newline(String message). - Implement branch-optimized zero-allocation
fastParseOn.
- Implement
-
Task 1.7: Epsilon Value Support & Primitives
- Extend
EpsilonParserto store an optional result valuevalue(defaulting tonull). - Provide a reusable static singleton
EpsilonParser.INSTANCEfornull. - Add static factories
Parser.epsilon()andParser.epsilon(Object value). - Add static factories
Parser.failure()andParser.failure(String message). - Add
Parser.constant(Object value)replacing results with a constant value.
- Extend
-
Task 1.8: Token & Context Formatting Utilities
- Add
Token.positionString(String buffer, int position)returning"line:column". - Add
Token.join(Iterable<Token> tokens)combining adjacent tokens from the same buffer. - Add
Context.toPositionString()returning"line:column".
- Add
-
Task 1.9: Phase 1 Parallel Unit Tests
- Create test classes under
src/test/java/org/petitparser/:org.petitparser.parser.actions.WhereParserTestorg.petitparser.parser.combinators.SkipParserTestorg.petitparser.parser.combinators.LabelParserTestorg.petitparser.parser.primitive.PositionParserTestorg.petitparser.parser.primitive.NewlineParserTestorg.petitparser.parser.primitive.EpsilonParserTestorg.petitparser.context.TokenTest(augment withjoinandpositionStringtests)
- Ensure 100% test coverage for all new lines and branches.
- Create test classes under
-
Task 2.1: RepeatingCharacterParser (Zero-List Lexing)
- Implement
org.petitparser.parser.repeating.RepeatingCharacterParserdirectly returning aStringviabuffer.substring(start, position)without intermediateList<Character>allocations. - Add
starString(),starString(String message)toParser. - Add
plusString(),plusString(String message)toParser. - Add
timesString(int count),timesString(int count, String message)toParser. - Add
repeatString(int min, int max),repeatString(int min, int max, String message)toParser. - Auto-specialize when called on
CharacterParserto instantiateRepeatingCharacterParser.
- Implement
-
Task 2.2: SeparatedList Data Structure
- Implement immutable
org.petitparser.parser.repeating.SeparatedList<R, S>holdingelementsandseparators. - Implement
getSequential()returning alternating interleaved elements and separators. - Implement functional folding:
foldLeft(FoldFunction<R, S> callback)andfoldRight(FoldFunction<R, S> callback). - Implement
equals,hashCode,toString, andIterable<Object>.
- Implement immutable
-
Task 2.3: SeparatedRepeatingParser
- Implement
org.petitparser.parser.repeating.SeparatedRepeatingParserimplementingSequentialParser. - Add
starSeparated(Parser separator)toParser. - Add
plusSeparated(Parser separator)toParser. - Add
timesSeparated(Parser separator, int count)toParser. - Add
repeatSeparated(Parser separator, int min, int max)toParser. - Retain existing
separatedByanddelimitedByreturningList<Object>for full backward compatibility.
- Implement
-
Task 2.4: Phase 2 Parallel Unit Tests
- Create test classes:
org.petitparser.parser.repeating.RepeatingCharacterParserTestorg.petitparser.parser.repeating.SeparatedListTestorg.petitparser.parser.repeating.SeparatedRepeatingParserTest
- Test zero, single, multiple repetitions, fast-parse, and folding operations.
- Create test classes:
-
Task 3.1: Dual-Nature Tuple Hierarchy (
Tuple2toTuple9)- Implement immutable
org.petitparser.utils.tuples.Tuple2<T1, T2>throughTuple9extendingjava.util.AbstractList<Object>. - Provide strongly-typed accessors:
tuple.first(),tuple.second(),tuple.third(), etc. - Implement
AbstractListcontract (size(),get(int index)), allowing tuples to be treated directly asList<Object>for 100% backward compatibility with legacy tests and downstream consumers (tuple.equals(Arrays.asList(...)) == true). - Implement zero-allocation accessors and structural equality (
equals,hashCode,toString).
- Implement immutable
-
Task 3.2: Multi-Arity Functional Interfaces
- Create standard functional interfaces in
org.petitparser.utils.functions:Function3<T1, T2, T3, R>,Function4<T1, T2, T3, T4, R>, ...Function9<...>.
- Ensure compatibility with Java
BiFunctionfor arity-2 sequences.
- Create standard functional interfaces in
-
Task 3.3: Typed Sequence Combinators (
SequenceParser2toSequenceParser9)- Implement
SequenceParser2<T1, T2>throughSequenceParser9<...>implementingSequentialParser. - Add fluent
.then(Parser<TN> next)onParserand sequence parsers, automatically flattening intoSequenceParser3,SequenceParser4, etc. up to 9 elements. - Add strongly-typed
.map(BiFunction<? super T1, ? super T2, ? extends R> function)onSequenceParser2. - Add strongly-typed
.map(Function3<...>)through.map(Function9<...>)onSequenceParser3throughSequenceParser9, eliminating the need for users to manually unpack tuple or list elements. - Implement zero-allocation
fastParseOn,copy, and equality methods.
- Implement
-
Task 3.4: Static Sequence Factories
- Add static factories
Parser.seq(Parser<T1> p1, Parser<T2> p2)returningSequenceParser2<T1, T2>. - Add static factories
Parser.seq(p1, p2, p3)throughParser.seq(p1, ..., p9)returning corresponding typed sequence parsers. - Retain existing
Parser.seq(Parser... parsers)and instance methodseq(Parser other)returningSequenceParser(Parser<List<Object>>) for backward compatibility.
- Add static factories
-
Task 3.5: Type Extraction, Casting & Permutation
- Implement
org.petitparser.parser.actions.CastParser<R, S>withcast()andcast(Class<S> clazz)onParser. - Implement
castList(Class<S> clazz)onParser<List<?>>. - Implement
pick(int index)with negative indexing support, specialized onRepeatingParser<E>to returnParser<E>. - Implement
permute(int... indices)returning reorderedTupleorList.
- Implement
-
Task 3.6: Variance-Friendly Alternatives (
or&orWiden)- Add
or(Parser<? extends R> other)toParser<R>to preserve precise local type inference when using Javavar. - Add
<O> Parser<O> orWiden(Parser<? extends O> other)and staticParser.or(Parser<? extends T>... parsers)for explicit widening across differing types (e.g.IntegerandDoubleintoNumber).
- Add
-
Task 3.7: Strongly-Typed Grammar Definitions
- Introduce
org.petitparser.tools.Production<T>typed token key:- Factory
Production.of(String name)/GrammarDefinition.production(String name).
- Factory
- Add typed
def(Production<T> production, Parser<? extends T> parser). - Add typed
ref(Production<T> production)returningParser<T>. - Retain string-based
def(String, Parser)and enhanceref(String)with<T> Parser<T> ref(String name)for seamless existing code compatibility.
- Introduce
-
Task 3.8: Phase 3 Parallel Unit Tests
- Create test classes:
org.petitparser.utils.tuples.TupleTest(verifying typed access andList<Object>contract)org.petitparser.parser.combinators.SequenceParserNTest(verifyingthen(),map(), arity 2-9, fast-parse)org.petitparser.parser.actions.CastParserTestorg.petitparser.parser.actions.PickParserTestorg.petitparser.tools.TypedGrammarDefinitionTest
- Create test classes:
-
Task 4.1: Indent Engine
- Implement
org.petitparser.tools.Indentclass managing an indentation stack and current indent string. - Implement
same()matching current indentation level. - Implement
increase()requiring deeper indentation and pushing to stack. - Implement
decrease()popping indentation from stack. - Implement
during(Parser inner)with transactional rollback on parse failure and choice backtracking.
- Implement
-
Task 4.2: Phase 4 Parallel Unit Tests
- Create
org.petitparser.tools.IndentTest:- Success restoring outer indentation.
- Failure rolling back indentation stack.
- Nested block scopes.
- Choice rollback trying alternate indented branches.
- Custom indentation token parsers (tabs vs spaces).
- Create
-
Task 5.1: Hybrid CharacterPredicate Hierarchy
- Preserve
@FunctionalInterface CharacterPredicateinterface (test(char)) for backward compatibility with user lambdas. - Introduce concrete AST classes implementing
CharacterPredicateandisEqualTo:-
SingleCharPredicate,RangeCharPredicate,RangesCharPredicate(binary search on primitive arrays). -
LookupCharPredicate($O(1)$ boolean array /BitSetlookup table for ASCII/Latin-1). -
NotCharPredicate(negated delegate). - Reusable singleton constants:
DigitCharPredicate,LetterCharPredicate,LowercaseCharPredicate,UppercaseCharPredicate,WhitespaceCharPredicate,WordCharPredicate,ConstantCharPredicate.any,ConstantCharPredicate.none.
-
- Update
CharacterParserfactories (digit(),letter(), etc.) to use AST singletons sop.isEqualTo(p)succeeds across separate calls.
- Preserve
-
Task 5.2: Range Merging & Character Optimization
- Implement
optimizedRanges(List<RangeCharPredicate> ranges):- Sort ranges by start and stop.
- Merge adjacent and overlapping ranges to minimize runtime branch tests.
- Choose optimal predicate (
$O(1)$ Lookup table, Single, Range, or Ranges with binary search).
- Implement
optimizedString(String string, boolean ignoreCase).
- Implement
-
Task 5.3: StringParser SIMD & Equality Fix
- Refactor
StringParser.ofIgnoringCaseto avoid non-comparable method reference lambdas (value::equalsIgnoreCase). - Implement
StringIgnoreCaseParserstoring the literal string and comparing literals inhasEqualProperties. - Use
String.regionMatchesandString.startsWithfor HotSpot-intrinsic performance.
- Refactor
-
Task 5.4: Unicode Code Points
- Implement
UnicodeCharacterParserdecoding UTF-16 surrogate pairs into 21-bit code points ($0 \dots \text{0x10FFFF}$ ). - Add
unicodeparameter / flag to character primitives (any,char,pattern,anyOf,noneOf).
- Implement
-
Task 5.5: Phase 5 Parallel Unit Tests
- Create tests:
org.petitparser.parser.primitive.CharacterPredicateAstTestorg.petitparser.parser.primitive.UnicodeCharacterParserTest- Verify structural equality of character parsers across all predicate variations.
- Create tests:
-
Task 6.1: ExpressionBuilder Enhancements
- Introduce generic typing
ExpressionBuilder<T>andExpressionGroup<T>. - Add
primitive(Parser parser)directly onExpressionBuilder. - Add
loopbackgetter onExpressionBuilder. - Add
optional(Object value)onExpressionGroup. - Refactor binary left/right operator builders to use
plusSeparatedandSeparatedList.foldLeft/foldRight. - Add typed functional callbacks (e.g. ternary
(left, op, right) -> resultviaFunction3) while preserving existingList<Object>callbacks for compatibility.
- Introduce generic typing
-
Task 6.2: Phase 6 Parallel Unit Tests
- Augment
org.petitparser.tools.ExpressionBuilderTestto test:- Builder-level primitives.
- Optional expression groups.
- Typed operator callbacks.
- Augment
-
Task 7.1: Grammar Graph Traversal & Analyzer Base
- Implement
org.petitparser.utils.Analyzerleveraging standard Java collections andStream:- Reachable parsers discovery (
parsers). - Deep children set cache (
allChildren(parser)). - Graph path search (
findPath,findPathTo,findAllPaths,findAllPathsTo).
- Reachable parsers discovery (
- Implement
-
Task 7.2: Grammar Property Computation
- Implement nullability fixed-point analysis (
isNullable(parser)). - Implement FIRST-set calculation (
firstSet(parser)): terminal parsers that can appear first, takingSequentialParserinto account. - Implement FOLLOW-set calculation (
followSet(parser)): terminal parsers that can immediately succeedparser. - Implement cycle detection (
cycleSet(parser)).
- Implement nullability fixed-point analysis (
-
Task 7.3: Grammar Reference Inlining
- Implement
resolve(Parser parser)resolving allResolvableParserreferences into direct cycles/graphs without delegate overhead.
- Implement
-
Task 7.4: Phase 7 Parallel Unit Tests
- Create
org.petitparser.utils.AnalyzerTest:- First-set and follow-set validation on LL/LR grammar definitions.
- Nullable chain detection.
- Path searching between arbitrary nodes.
- Resolvable reference inlining.
- Create
-
Task 8.1: Linter Engine Architecture
- Implement
org.petitparser.utils.linter.LinterRule,LinterIssue,LinterType(info, warning, error). - Implement
linter(Parser root, ...)executing active rules usingAnalyzer.
- Implement
-
Task 8.2: 13 Linter Rules Implementation
-
CharacterRepeaterRule: Identifies.star().flatten()on character parsers and suggestsstarString(). -
DuplicateParserRule: Identifies duplicate structurally equal parser instances in grammar graph. -
LeftRecursionRule: Identifies left-recursive loops that lead to infinite recursion. -
NestedChoiceRule: Identifies nested choice combinators. -
NullableRepeaterRule: Identifies repeating parsers over nullable delegates that lead to infinite loops. -
OverlappingChoiceRule: Identifies choices where a prefix choice shadows a later choice. -
RepeatedChoiceRule: Identifies duplicate branches in a choice parser. -
UnnecessaryFlattenRule: Identifiesflatten()on parsers already producing Strings. -
UnnecessaryResolvableRule: Identifies unresolved resolvable wrappers. -
UnoptimizedFlattenRule: Identifiesflatten()without error messages preventing switch to fast parsing mode. -
UnreachableChoiceRule: Identifies choices located after unconditional matchers. -
UnresolvedSettableRule: IdentifiesSettableParserleft in undefined state. -
UnusedResultRule: Identifies complex sub-parses whose results are discarded (e.g. inner actions insideflatten()).
-
-
Task 8.3: Phase 8 Parallel Unit Tests
- Create
org.petitparser.utils.linter.LinterTestverifying positive and negative triggers for each of the 13 rules.
- Create
-
Task 9.1: Extensible OptimizeRule Framework
- Define
OptimizeRuleinterface and refactorOptimizerto use modular rules. - Keep
RemoveDelegateandRemoveDuplicaterules.
- Define
-
Task 9.2: New Optimizer Rules
- Implement
FlattenChoiceRule: Flattens nested choices[a, [b, c]]into[a, b, c]. - Implement
CharacterRepeaterRule: TransformsFlattenParser(PossessiveRepeatingParser(CharacterParser))intoRepeatingCharacterParser.
- Implement
-
Task 9.3: Phase 9 Parallel Unit Tests
- Augment
org.petitparser.utils.OptimizerTestto cover choice flattening and character repeater rewrites.
- Augment
-
Task 10.1: Progress Stepper
- Implement
org.petitparser.utils.Progress:progress(Parser root, Consumer<ProgressFrame> observer)visual execution debugger.ProgressFramecapturing parser, context, position, and backtracking events.
- Implement
-
Task 10.2: Matcher Offset & Lazy Streams
- Enhance
accept(String input, int start)with start offset. - Add lazy
matchesAsStream(String input, int start)returningjava.util.stream.Stream<T>backed by a customSpliterator. - Add lazy
matches(String input, int start)returningIterable<T>.
- Enhance
-
Task 10.3: Phase 10 Parallel Unit Tests
- Create
org.petitparser.utils.ProgressTestandorg.petitparser.MatcherTest.
- Create
-
Task 11.1: Downstream Module Compatibility Verification
- Verify
petitparser-json,petitparser-xml, andpetitparser-smalltalkcompile and pass tests without modifications. - Ensure raw-type usage compiles cleanly without errors or breaking changes for existing code.
- Verify
-
Task 11.2: JMH Benchmark Suite
- Measure throughput and allocation rate of
RepeatingCharacterParservs.star().flatten(). - Measure
$O(1)$ Lookup table predicates vs chained range predicates. - Measure
fastParseOnzero-allocation performance against Dart baseline.
- Measure throughput and allocation rate of