Quickly compare a string against a Collection in Java

Question

I am trying to calculate edit distances of a string against a collection to find the closest match. My current problem is that the collection is very large (about 25000 items), so I had to narrow down the set to just strings of similar lengths but that still would only narrow it down to a few thousand strings and this still is very slow. Is there a datastructure that allows for a quick lookup of similar strings or is there another way I could address this problem?

By similar I mean comparing words that are common spelling mistakes such as "exanple" and "example" or "weird" and "wierd". — Lezan, Feb 04 '12 at 09:01
Looks like you want an implementation of the levenstein distance: http://stackoverflow.com/questions/6087281/similarity-score-levenshtein. — Kurt Du Bois, Feb 04 '12 at 09:05
I am currently doing it the following way: String currentString; List distanceList; (for word: wordList){ int distance = calculateDistance(currentString,word) distanceList.add(distance) } — Lezan, Feb 04 '12 at 09:07

score 8 · Accepted Answer · answered Feb 04 '12 at 08:50

8

Sounds like a BK-tree might be what you want. Here's an article discussing them: http://blog.notdot.net/2007/4/Damn-Cool-Algorithms-Part-1-BK-Trees. A quick Google yields some Java implementations.

answered Feb 04 '12 at 08:50

SimonC

6,590
1
23
40

Thanks I will look this up and let you know how it goes, thank you! – Lezan Feb 04 '12 at 09:07
Yup that did it, needed a different implementation of the search but it was perfect! Thank you!! – Lezan Feb 05 '12 at 11:16

score 6 · Answer 2 · answered Feb 04 '12 at 10:32

6

Levenshtein Automata allow for fast selection of a set of words from a large dictionary such that they are within the given Levenshtein distance from a given word.

See: Schulz K, Mihov S. (2002) Fast String Correction with Levenshtein-Automata.

answered Feb 04 '12 at 10:32

kkm inactive - support strike

5,190
2
32
59

score 2 · Answer 3 · answered Feb 04 '12 at 08:42

2

If your criteria for 'similar' define a total ordering, you should be able to define a Comparator and use a TreeSet to find the closest matches (eg using the ceiling and floor methods).

answered Feb 04 '12 at 08:42

Mark Rotteveel

100,966
191
140
197

Quickly compare a string against a Collection in Java

3 Answers3

Linked