Skip to content

Charset in meta content does not correctly parse for trailing semi-colon #92

Description

@1619digital

Reference: http://www.w3.org/html/wg/drafts/html/master/infrastructure.html#algorithm-for-extracting-a-character-encoding-from-a-meta-element

Because the ContentAttrParser is looking only for a space character to terminate an unquoted charset

<meta http-equiv="Content-Type" content="charset=iso8859-2;text/html">

will incorrectly be inferred to have the charset 'iso8859-2;text/html'. The fix is to add a semicolon to the spaceCharacters scanned in SkipUntil - line 860.

EDIT: as per specification. Also, I don't know what the status is of the parser tests, but they're out of date and incorrect and (obviously) not used. Although most of the tests are still valid, so it would not take much to bring them back into the full test regime.

Activity

  1. modified the milestones: 1.0, on Jun 6, 2016
  2. modified the milestones: , 1.0 on Oct 3, 2017
  3. removed this from the 1.0 milestone on Oct 31, 2017
  4. willkg commented on Oct 31, 2017

    @willkg
    Contributor

    If someone needs this, please submit a PR.

  5. gsnedders commented on Nov 8, 2017

    @gsnedders
    Member

    We have a problem with the tests for this that they currently don't agree whether they're testing the eventual encoding (including potentially with the tokenizer changing the encoding while parsing, after the pre-parse) or just the pre-parse.

    See html5lib/html5lib-tests#28 for that.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions