Fuzzy Regular Expressions
Variants of regular expressions can be used for working with text in natural language, when it is necessary to take into account possible typos and spelling variants. For example, the text "Julius Caesar" might be a fuzzy match for:
- Gaius Julius Caesar
- Yulius Cesar
- G. Juliy Caezar
In such cases the mechanism implements some fuzzy string matching algorithm and possibly some algorithm for finding the similarity between text fragment and pattern.
This task is closely related to both full text search and named entity recognition.
Some software libraries work with fuzzy regular expressions:
- TRE - well-developed portable free project in C, which uses syntax similar to POSIX
- FREJ - open source project in Java with non-standard syntax (which utilizes prefix, Lisp-like notation), targeted to allow easy use of substitutions of inner matched fragments in outer blocks, but lacks many features of standard regular expressions.
- agrep - command-line utility (proprietary, but free for non-commercial usage).
Read more about this topic: Regular Expression
Famous quotes containing the words fuzzy, regular and/or expressions:
“Even their song is not a sure thing.
It is not a language;
it is a kind of breathing.
They are two asthmatics
whose breath sobs in and out
through a small fuzzy pipe.”
—Anne Sexton (19281974)
“The solid and well-defined fir-tops, like sharp and regular spearheads, black against the sky, gave a peculiar, dark, and sombre look to the forest.”
—Henry David Thoreau (18171862)
“Its idea of production value is spending a million dollars dressing up a story that any good writer would throw away. Its vision of the rewarding movie is a vehicle for some glamour-puss with two expressions and eighteen changes of costume, or for some male idol of the muddled millions with a permanent hangover, six worn-out acting tricks, the build of a lifeguard, and the mentality of a chicken-strangler.”
—Raymond Chandler (18881959)