在 Java 中,可以用正则表达式的“反向引用(backreference)”来匹配重复出现的词。下面按常见场景说明。
import java.util.regex.*;
public class Test {
public static void main(String[] args) {
String text = "hello hello world java java";
Pattern p = Pattern.compile("\\b(\\w+)\\s+\\1\\b");
Matcher m = p.matcher(text);
while (m.find()) {
System.out.println(m.group());
}
}
}
hello hello
java java
\\b:单词边界(\\w+):捕获一个单词\\s+:空白字符\\1:反向引用,表示和前面捕获的单词完全相同\\b:单词结束✅ 可以匹配 连续重复的词
Pattern p = Pattern.compile("\\b(\\w+)\\s+\\1\\b", Pattern.CASE_INSENSITIVE);
这样 Hello hello 也能匹配。
如果你想找出在文本中重复出现过的词(不一定相邻):
import java.util.*;
import java.util.regex.*;
public class Test {
public static void main(String[] args) {
String text = "java is good and java is powerful";
Map<String, Integer> map = new HashMap<>();
Pattern p = Pattern.compile("\\b\\w+\\b");
Matcher m = p.matcher(text);
while (m.find()) {
String word = m.group().toLowerCase();
map.put(word, map.getOrDefault(word, 0) + 1);
}
map.forEach((k, v) -> {
if (v > 1) {
System.out.println(k + " 重复了 " + v + " 次");
}
});
}
}
Pattern p = Pattern.compile("(\\p{L}+)\\s+\\1");
\\p{L}:匹配任意语言的字母(包括中文)示例:
String text = "你好 你好 世界";
❌ 错误写法:
"\\b(\\w+)\\s+\\1"
✅ 正确写法(Java 字符串中):
"\\b(\\w+)\\s+\\1\\b"
因为 \1 在字符串里要写成 \\1
| 需求 | 正则 |
|---|---|
| 连续重复单词 | \\b(\\w+)\\s+\\1\\b |
| 忽略大小写 | Pattern.CASE_INSENSITIVE |
| 中文重复词 | (\\p{L}+)\\s+\\1 |
| 全文统计重复 | 正则 + Map |
如果你有具体文本示例(比如日志、句子、特殊分隔符),我可以直接帮你写完整正则。
免责声明:本站发布的内容(图片、视频和文字)以原创、转载和分享为主,文章观点不代表本网站立场,如果涉及侵权请联系站长邮箱:is@yisu.com进行举报,并提供相关证据,一经查实,将立刻删除涉嫌侵权内容。