Module Regexp.PCRE2.OPTION

Description

contains all option constants


Constant ALLOW_EMPTY_CLASS

optional constant Regexp.PCRE2.OPTION.ALLOW_EMPTY_CLASS

Description

(from the pcre2api manpage) By default, for compatibility with Perl, a closing square bracket thatimmediately follows an opening one is treated as a data character for the class. When ALLOW_EMPTY_CLASS is set, it terminates the class, which therefore contains no characters and so can never match.

Note

This symbol is not present in all versions of PCRE2. It was added in PCRE2 10.30.


Constant ALT_BSUX

optional constant Regexp.PCRE2.OPTION.ALT_BSUX

Description

(from the pcre2api manpage) This option request alternative handling of three escape sequences, which makes PCRE2's behaviour more like ECMAscript (aka JavaScript). When it is set:

  1. \U matches an upper case "U" character; by default \U causes a compile time error (Perl uses \U to upper case subsequent characters).

  2. \u matches a lower case "u" character unless it is followed by four hexadecimal digits, in which case the hexadecimal number defines the code point to match. By default, \u causes a compile time error (Perl uses it to upper case the following character).

  3. \x matches a lower case "x" character unless it is followed by two hexadecimal digits, in which case the hexadecimal number defines the code point to match. By default, as in Perl, a hexadecimal number is always expected after \x, but it may have zero, one, or two digits (so, for example, \xz matches a binary zero character followed by z).

Note

This symbol is not present in all versions of PCRE2. It was added in PCRE2 10.32.


Constant ALT_CIRCUMFLEX

constant Regexp.PCRE2.OPTION.ALT_CIRCUMFLEX

Description

(from the pcre2api manpage) In multiline mode (when MULTILINE is set), the circumflex metacharacter matches at the start of the subject, and also after any internal newline. However, it does not match after a newline at the end of the subject, for compatibility with Perl. If you want a multiline circumflex also to match after a terminating newline, you must set ALT_CIRCUMFLEX.


Constant ALT_EXTENDED_CLASS

optional constant Regexp.PCRE2.OPTION.ALT_EXTENDED_CLASS

Description

(from the pcre2api manpage) Alters the parsing of character classes to follow the extended syntax described by Unicode UTS#18. The ALT_EXTENDED_CLASS option has no impact on the behaviour of the Perl-specific "(?[...])" syntax for extended classes, but instead enables the alternative syntax of extended class behaviour inside ordinary "[...]" character classes.

Note

This symbol is not present in all versions of PCRE2. It was added in PCRE2 10.45.


Constant ALT_VERBNAMES

constant Regexp.PCRE2.OPTION.ALT_VERBNAMES

Description

(from the pcre2api manpage) By default, for compatibility with Perl, the name in any verb sequence such as (*MARK:NAME) is any sequence of characters that does not include a closing parenthesis. The name is not processed in any way, and it is not possible to include a closing parenthesis in the name. However, if the ALT_VERBNAMES option is set, normal backslash processing is applied to verb names and only an unescaped closing parenthesis terminates the name. A closing parenthesis can be included in a name either as \) or between \Q and \E. If the EXTENDED or EXTENDED_MORE option is set with ALT_VERBNAMES, unescaped whitespace in verb names is skipped and #-comments are recognized, exactly as in the rest of the pattern.


Constant ANCHORED

constant Regexp.PCRE2.OPTION.ANCHORED

Description

(from the pcre2api manpage) If this bit is set, the pattern is forced to be "anchored", that is, it is constrained to match only at the first matching point in the string that is being searched (the "subject string"). This effect can also be achieved by appropriate constructs in the pattern itself, which is the only way to do it in Perl.


Constant AUTO_CALLOUT

constant Regexp.PCRE2.OPTION.AUTO_CALLOUT

Description

(from the pcre2api manpage) If this bit is set, callout items are automatically inserted, all with number 255, before each pattern item, except immediately before or after an explicit callout in the pattern.


Constant CASELESS

constant Regexp.PCRE2.OPTION.CASELESS

Description

(from the pcre2api manpage) If this bit is set, letters in the pattern match both upper and lower case letters in the subject. It is equivalent to Perl's /i option, and it can be changed within a pattern by a (?i) option setting


Constant DOLLAR_ENDONLY

constant Regexp.PCRE2.OPTION.DOLLAR_ENDONLY

Description

(from the pcre2api manpage) If this bit is set, a dollar metacharacter in the pattern matches only at the end of the subject string. Without this option, a dollar also matches immediately before a newline at the end of the string (but not before any other newlines). The DOLLAR_ENDONLY option is ignored if MULTILINE is set. There is no equivalent to this option in Perl, and no way to set it within a pattern.


Constant DOTALL

constant Regexp.PCRE2.OPTION.DOTALL

Description

(from the pcre2api manpage) If this bit is set, a dot metacharacter in the pattern matches any character, including one that indicates a newline. However, it only ever matches one character, even if newlines are coded as CRLF. Without this option, a dot does not match when the current position in the subject is at a newline. This option is equivalent to Perl's /s option, and it can be changed within a pattern by a (?s) option setting. A negative class such as [^a] always matches newline characters, and the \N escape sequence always matches a non-newline character, independent of the setting of DOTALL.


Constant DUPNAMES

constant Regexp.PCRE2.OPTION.DUPNAMES

Description

(from the pcre2api manpage) If this bit is set, names used to identify capture groups need not be unique. This can be helpful for certain types of pattern when it is known that only one instance of the named group can ever be matched.


Constant ENDANCHORED

constant Regexp.PCRE2.OPTION.ENDANCHORED

Description

(from the pcre2api manpage) If this bit is set, the end of any pattern match must be right at the end of the string being searched (the "subject string"). If the pattern match succeeds by reaching (*ACCEPT), but does not reach the end of the subject, the match fails at the current starting point. For unanchored patterns, a new match is then tried at the next starting point. However, if the match succeeds by reaching the end of the pattern, but not the end of the subject, backtracking occurs and an alternative match may be found.


Constant EXTENDED

constant Regexp.PCRE2.OPTION.EXTENDED

Description

(from the pcre2api manpage) If this bit is set, most white space characters in the pattern are totally ignored except when escaped, inside a character class, or inside a \Q...\E sequence. However, white space is not allowed within sequences such as (?> that introduce various parenthesized groups, nor within numerical quantifiers such as {1,3}. Ignorable white space is permitted between an item and a following quantifier and between a quantifier and a following + that indicates possessiveness. EXTENDED is equivalent to Perl's /x option, and it can be changed within a pattern by a (?x) option setting.


Constant EXTENDED_MORE

constant Regexp.PCRE2.OPTION.EXTENDED_MORE

Description

(from the pcre2api manpage) This option has the effect of EXTENDED, but, in addition, unescaped space and horizontal tab characters are ignored inside a character class. Note: only these two characters are ignored, not the full set of pattern white space characters that are ignored outside a character class. EXTENDED_MORE is equivalent to Perl's /xx option, and it can be changed within a pattern by a (?xx) option setting.


Constant FIRSTLINE

constant Regexp.PCRE2.OPTION.FIRSTLINE

Description

(from the pcre2api manpage) If this option is set, the start of an unanchored pattern match must be before or at the first newline in the subject string following the start of matching, though the matched text may continue over the newline. If startoffset is non-zero, the limiting newline is not necessarily the first newline in the subject. For example, if the subject string is "abc\nxyz" (where \n represents a single-character newline) a pattern match for "yz" succeeds with FIRSTLINE if startoffset is greater than 3.


Constant LITERAL

optional constant Regexp.PCRE2.OPTION.LITERAL

Description

(from the pcre2api manpage) If this option is set, all meta-characters in the pattern are disabled, and it is treated as a literal string. Matching literal strings with a regular expression engine is not the most efficient way of doing it. If you are doing a lot of literal matching and are worried about efficiency, you should consider using other approaches.

Note

This symbol is not present in all versions of PCRE2. It was added in PCRE2 10.30.


Constant MATCH_INVALID_UTF

optional constant Regexp.PCRE2.OPTION.MATCH_INVALID_UTF

Description

(from the pcre2api manpage) This option forces UTF (see below) and also enables support for matching in subject strings that contain invalid UTF sequences.

Note

This symbol is not present in all versions of PCRE2. It was added in PCRE2 10.34.


Constant MATCH_UNSET_BACKREF

constant Regexp.PCRE2.OPTION.MATCH_UNSET_BACKREF

Description

(from the pcre2api manpage) If this option is set, a backreference to an unset capture group matches an empty string (by default this causes the current matching alternative to fail). A pattern such as (\1)(a) succeeds when this option is set (assuming it can find an "a" in the subject), whereas it fails by default, for Perl compatibility. Setting this option makes PCRE2 behave more like ECMAscript (aka JavaScript).


Constant MULTILINE

constant Regexp.PCRE2.OPTION.MULTILINE

Description

(from the pcre2api manpage) By default, for the purposes of matching "start of line" and "end of line", PCRE2 treats the subject string as consisting of a single line of characters, even if it actually contains newlines. The "start of line" metacharacter (^) matches only at the start of the string, and the "end of line" metacharacter ($) matches only at the end of the string, or before a terminating newline (except when DOLLAR_ENDONLY is set). Note, however, that unless DOTALL is set, the "any character" metacharacter (.) does not match at a newline. This behaviour (for ^, $, and dot) is the same as Perl.

When MULTILINE it is set, the "start of line" and "end of line" constructs match immediately following or immediately before internal newlines in the subject string, respectively, as well as at the very start and end. This is equivalent to Perl's /m option, and it can be changed within a pattern by a(?m) option setting. Note that the "start of line" metacharacter does not match after a newline at the end of the subject, for compatibility with Perl. However, you can change this by setting the ALT_CIRCUMFLEX option. If there are no newlines in a subject string, or no occurrences of ^ or $ in a pattern, setting MULTILINE has no effect.


Constant NEVER_BACKSLASH_C

constant Regexp.PCRE2.OPTION.NEVER_BACKSLASH_C

Description

(from the pcre2api manpage) This option locks out the use of \C in the pattern that is being compiled. This escape can cause unpredictable behaviour in UTF-8 or UTF-16 modes, because it may leave the current matching point in the middle of a multi-code-unit character. This option may be useful in applications that process patterns from external sources. Note that there is also a build-time option that permanently locks out the use of \C.


Constant NEVER_UCP

constant Regexp.PCRE2.OPTION.NEVER_UCP

Description

(from the pcre2api manpage) This option locks out the use of Unicode properties for handling \B, \b, \D, \d, \S, \s, \W, \w, and some of the POSIX character classes, as described for the UCP option below. In particular, it prevents the creator of the pattern from enabling this facility by starting the pattern with (*UCP). This option may be useful in applications that process patterns from external sources. The option combination UCP and NEVER_UCP causes an error.


Constant NEVER_UTF

constant Regexp.PCRE2.OPTION.NEVER_UTF

Description

(from the pcre2api manpage) This option locks out interpretation of the pattern as UTF-8, UTF-16, or UTF-32. In particular, it prevents the creator of the pattern from switching to UTF interpretation by starting the pattern with (*UTF). This option may be useful in applications that process patterns from external sources. The combination of UTF and NEVER_UTF causes an error.


Constant NO_AUTO_CAPTURE

constant Regexp.PCRE2.OPTION.NO_AUTO_CAPTURE

Description

(from the pcre2api manpage) If this option is set, it disables the use of numbered capturing parentheses in the pattern. Any opening parenthesis that is not followed by ? behaves as if it were followed by ?: but named parentheses can still be used for capturing (and they acquire numbers in the usual way). This is the same as Perl's /n option. Note that, when this option is set, references to capture groups (backreferences or recursion/subroutine calls) may only refer to named groups, though the reference can be by name or by number.


Constant NO_AUTO_POSSESS

constant Regexp.PCRE2.OPTION.NO_AUTO_POSSESS

Description

(from the pcre2api manpage) If this (deprecated) option is set, it disables "auto-possessification", which is an optimization that, for example, turns a+b into a++b in order to avoid backtracks into a+ that can never be successful. However, if callouts are in use, auto-possessification means that some callouts are never taken. You can set this option if you want the matching functions to do a full unoptimized search and run all the callouts, but it is mainly provided for testing purposes.


Constant NO_DOTSTAR_ANCHOR

constant Regexp.PCRE2.OPTION.NO_DOTSTAR_ANCHOR

Description

(from the pcre2api manpage) If this (deprecated) option is set, it disables an optimization that is applied when .* is the first significant item in a top-level branch of a pattern, and all the other branches also start with .* or with \A or \G or ^. The optimization is automatically disabled for .* if it is inside an atomic group or a capture group that is the subject of a backreference, or if the pattern contains (*PRUNE) or (*SKIP). When the optimization is not disabled, such a pattern is automatically anchored if DOTALL is set for all the .* items and MULTILINE is not set for any ^ items. Otherwise, the fact that any match must start either at the start of the subject or following a newline is remembered. Like other optimizations, this can cause callouts to be skipped.


Constant NO_START_OPTIMIZE

constant Regexp.PCRE2.OPTION.NO_START_OPTIMIZE

Description

(from the pcre2api manpage) This is an option whose main effect is at matching time.

There are a number of optimizations that may occur at the start of a match, in order to speed up the process. For example, if it is known that an unanchored match must start with a specific code unit value, the matching code searches the subject for that value, and fails immediately if it cannot find it, without actually running the main matching function. The start-up optimizations are in effect a pre-scan of the subject that takes place before the pattern is run.

Disabling the start-up optimizations may cause performance to suffer. However, this may be desirable for patterns which contain callouts or items such as (*COMMIT) and (*MARK).


Constant NO_UTF_CHECK

constant Regexp.PCRE2.OPTION.NO_UTF_CHECK

Description

(from the pcre2api manpage) When UTF is set, the validity of the pattern as a UTF string is automatically checked. If an invalid UTF sequence is found, an error is thrown.

If you know that your pattern is a valid UTF string, and you want to skip this check for performance reasons, you can set the NO_UTF_CHECK option. When it is set, the effect of passing an invalid UTF string a s a pattern is undefined. It may cause your program to crash or loop.

Note also that setting NO_UTF_CHECK at compile time does not disable the error that is given if an escape sequence for an invalid Unicode code point is encountered in the pattern. In particular, the so-called "surrogate" code points (0xd800 to 0xdfff) are invalid.


Constant UCP

constant Regexp.PCRE2.OPTION.UCP

Description

(from the pcre2api manpage) This option has two effects. Firstly, it change the way PCRE2 processes \B, \b, \D, \d, \S, \s, \W, \w, and some of the POSIX character classes. By default, only ASCII characters are recognized, but if UCP is set,' Unicode properties are used to classify characters.

The second effect of UCP is to force the use of Unicode properties for upper/lower casing operations, even when UTF is not set. This makes it possible to process strings in the 16-bit UCS-2 code.

In Pike the option is automatically enabled unless either UTF or NEVER_UCP is set.


Constant UNGREEDY

constant Regexp.PCRE2.OPTION.UNGREEDY

Description

(from the pcre2api manpage) This option inverts the "greediness" of the quantifiers so that they are not greedy by default, but become greedy if followed by "?". It is not compatible with Perl. It can also be set by a (?U) option setting within the pattern.


Constant UTF

constant Regexp.PCRE2.OPTION.UTF

Description

(from the pcre2api manpage) This option causes PCRE2 to regard both the pattern and the subject strings that are subsequently processed as strings of UTF characters instead of single-code-unit strings.