How to use a Lucene Analyzer to tokenize a String?

Question

Is there a simple way I could use any subclass of Lucene's Analyzer to parse/tokenize a String?

Something like:

String to_be_parsed = "car window seven";
Analyzer analyzer = new StandardAnalyzer(...);
List<String> tokenized_string = analyzer.analyze(to_be_parsed);

That's a pretty vague question you're asking. The answer is "Yes". But it depends a lot on *how* you want to parse/tokenize said string. — stevevls, Jun 13 '11 at 18:39
@stevevls added an example. I used List but it doesn't have to be necessarly a List. — Felipe Hummel, Jun 13 '11 at 18:44

score 57 · Answer 1 · edited Sep 25 '13 at 12:27

57

Based off of the answer above, this is slightly modified to work with Lucene 4.0.

public final class LuceneUtil {

  private LuceneUtil() {}

  public static List<String> tokenizeString(Analyzer analyzer, String string) {
    List<String> result = new ArrayList<String>();
    try {
      TokenStream stream  = analyzer.tokenStream(null, new StringReader(string));
      stream.reset();
      while (stream.incrementToken()) {
        result.add(stream.getAttribute(CharTermAttribute.class).toString());
      }
    } catch (IOException e) {
      // not thrown b/c we're using a string reader...
      throw new RuntimeException(e);
    }
    return result;
  }

}

edited Sep 25 '13 at 12:27

Parvin Gasimzade

25,180
8
56
83

answered Mar 05 '12 at 07:03

Ben McCann

18,548
25
83
101

13

In Lucene 4.1 you also need to add `stream.reset()` before the `while` statement – prestomanifesto Feb 28 '13 at 19:29
2

You may want to add a `stream.end(); stream.close();` after the while slope. – membersound Sep 04 '14 at 07:42
Note; above works perfectly in Lucene **7.0.1**. Just add sugar with `try-with-resource` on `TokenStream`. – earcam Oct 08 '17 at 23:27

stevevls · Accepted Answer · 2011-06-13T19:20:50.957

As far as I know, you have to write the loop yourself. Something like this (taken straight from my source tree):

public final class LuceneUtils {

    public static List<String> parseKeywords(Analyzer analyzer, String field, String keywords) {

        List<String> result = new ArrayList<String>();
        TokenStream stream  = analyzer.tokenStream(field, new StringReader(keywords));

        try {
            while(stream.incrementToken()) {
                result.add(stream.getAttribute(TermAttribute.class).term());
            }
        }
        catch(IOException e) {
            // not thrown b/c we're using a string reader...
        }

        return result;
    }  
}

Just one more note: As of Lucene 3.2 TermAttribute is deprecated in favor of CharTermAttribute. — Felipe Hummel, Jun 13 '11 at 19:30

score 2 · Answer 3 · answered Feb 07 '19 at 17:40

Even better by using try-with-resources! This way you don't have to explicitly call .close() that is required in higher versions of the library.

public static List<String> tokenizeString(Analyzer analyzer, String string) {
  List<String> tokens = new ArrayList<>();
  try (TokenStream tokenStream  = analyzer.tokenStream(null, new StringReader(string))) {
    tokenStream.reset();  // required
    while (tokenStream.incrementToken()) {
      tokens.add(tokenStream.getAttribute(CharTermAttribute.class).toString());
    }
  } catch (IOException e) {
    new RuntimeException(e);  // Shouldn't happen...
  }
  return tokens;
}

And the Tokenizer version:

  try (Tokenizer standardTokenizer = new HMMChineseTokenizer()) {
    standardTokenizer.setReader(new StringReader("我说汉语说得很好"));
    standardTokenizer.reset();
    while(standardTokenizer.incrementToken()) {
      standardTokenizer.getAttribute(CharTermAttribute.class).toString());
    }
  } catch (IOException e) {
      new RuntimeException(e);  // Shouldn't happen...
  }

score 2 · Answer 4 · answered Sep 20 '20 at 19:05

The latest best practices, as another Stack Overflow answer indicates, seems to be to add an attribute to the token stream and later access that attribute, rather than getting an attribute directly from the token stream. And for good measure you can make sure the analyzer gets closed. Using the very latest Lucene (currently v8.6.2) the code would look like this:

String text = "foo bar";
String fieldName = "myField";
List<String> tokens = new ArrayList();
try (Analyzer analyzer = new StandardAnalyzer()) {
  try (final TokenStream tokenStream = analyzer.tokenStream(fieldName, text)) {
    CharTermAttribute charTermAttribute = tokenStream.addAttribute(CharTermAttribute.class);
    tokenStream.reset();
    while(tokenStream.incrementToken()) {
      tokens.add(charTermAttribute.toString());
    }
    tokenStream.end();
  }
}

After that code is finished, tokens will contain a list of parsed tokens.

How to use a Lucene Analyzer to tokenize a String?

4 Answers4

Linked