Hoogle Search
Within LTS Haskell 24.62 (ghc-9.10.3)
Note that Stackage only displays results for the latest LTS and Nightly snapshot. Learn more.
validateUtf8More :: Utf8State -> ByteString -> (Int, Maybe Utf8State)text Data.Text.Internal.Encoding Validate another ByteString chunk in an ongoing stream of UTF-8-encoded text. Returns a pair:
- The first component n is the end position, relative to the current chunk, of the longest prefix of the accumulated bytestring which is valid UTF-8. n may be negative: that happens when an incomplete code point started in a previous chunk and is not completed by the current chunk (either that code point is still incomplete, or it is broken by an invalid byte).
- The second component ms indicates the
following:
- if ms = Nothing, the remainder of the chunk contains an invalid byte, within four bytes from position n;
- if ms = Just s', you can carry on validating another chunk by calling validateUtf8More with the new state s'.
Properties
Given:validateUtf8More s chunk = (n, ms)
- If the chunk is invalid, it cannot be extended to be valid.
ms = Nothing ==> validateUtf8More s (chunk <> more) = (n, Nothing)
- Validating two chunks sequentially is the same as validating them
together at once:
ms = Just s' ==> validateUtf8More s (chunk <> more) = first (length chunk +) (validateUtf8More s' more)
-
text Data.Text.Internal.Encoding.Utf16 No documentation available.
validate2 :: Word16 -> Word16 -> Booltext Data.Text.Internal.Encoding.Utf16 No documentation available.
-
text Data.Text.Internal.Encoding.Utf32 No documentation available.
-
text Data.Text.Internal.Encoding.Utf8 No documentation available.
validate2 :: Word8 -> Word8 -> Booltext Data.Text.Internal.Encoding.Utf8 No documentation available.
validate3 :: Word8 -> Word8 -> Word8 -> Booltext Data.Text.Internal.Encoding.Utf8 No documentation available.
validate4 :: Word8 -> Word8 -> Word8 -> Word8 -> Booltext Data.Text.Internal.Encoding.Utf8 No documentation available.
module Data.Text.Internal.
Validate Test whether or not a sequence of bytes is a valid UTF-8 byte sequence. In the GHC Haskell ecosystem, there are several representations of byte sequences. The only one that the stable text API concerns itself with is ByteString. Part of bytestring-to-text decoding is isValidUtf8ByteString, a high-performance UTF-8 validation routine written in C++ with fallbacks for various platforms. The C++ code backing this routine is nontrivial, so in the interest of reuse, this module additionally exports functions for working with the GC-managed ByteArray type. These ByteArray functions are not used anywhere else in text. They are for the benefit of library and application authors who do not use ByteString but still need to interoperate with text.
isValidUtf8ByteArray :: ByteArray -> Int -> Int -> Booltext Data.Text.Internal.Validate For pinned byte arrays larger than 128KiB, this switches to the safe FFI so that it does not prevent GC. This threshold (128KiB) was chosen somewhat arbitrarily and may change in the future.