gel*_*d0r 3 java eclipse encoding aspectj gradle
在使用 Java 的相等性检查(直接或间接)时,我遇到了德语“Umlaute”(ä、ö、ü、ß)的奇怪行为。从 Eclipse 运行、调试或测试时,一切都按预期工作,并且包含“Umlaute”的输入是被视为平等或不符合预期。
但是,当我使用 Spring Boot 构建应用程序并运行它时,对于包含“Umlaute”的单词(即像“Nationalität”这样的单词),这些相等性检查将失败。
通过 Jsoup 从网页中检索输入,并为某些关键字提取表格内容。页面的编码是 UTF-8,如果不是这种情况,我已经为 Jsoup 进行了处理以将其转换。源文件的编码也是 UTF-8。
Connection connection = Jsoup.connect(url)
.header("accept-language", "de-de, de, en")
.userAgent("Mozilla/5.0")
.timeout(10000)
.method(Method.GET);
Response response = connection.execute();
if(logger.isDebugEnabled())
logger.debug("Encoding of response: " +response.charset());
Document doc;
if(response.charset().equalsIgnoreCase("UTF-8"))
{
logger.debug("Response has expected charset");
doc = Jsoup.parse(response.body(), baseURL);
}
else
{
logger.debug("Response doesn't have exepcted charset and is converted");
doc = Jsoup.parse(new String(response.bodyAsBytes(), "UTF-8"), baseURL);
}
logger.debug("Encoding of document: " +doc.charset());
if(!doc.charset().equals(Charset.forName("UTF-8")))
{
logger.debug("Changing encoding of document from " +doc.charset());
doc.updateMetaCharsetElement(true);
doc.charset(Charset.forName("UTF-8"));
logger.debug("Changed encoding of document to: " +doc.charset());
}
return doc;
Run Code Online (Sandbox Code Playgroud)
阅读内容的示例日志输出(来自已部署的应用程序)。
Encoding of response: utf-8
Response has expected charset
Encoding of document: UTF-8
Run Code Online (Sandbox Code Playgroud)
示例输入:
<tr><th>Nationalität:</th> <td> [...] </td> </tr>
Run Code Online (Sandbox Code Playgroud)
包含 ä、ö、ü 或 ß 的单词失败但适用于其他单词的示例代码:
Element header = row.select("th").first();
String text = header.ownText();
if("Nationalität:".equals(text))
{
// goes here in eclipse
}
else
{
// and here in deployed spring boot app
}
Run Code Online (Sandbox Code Playgroud)
从 Eclipse 运行和我缺少的构建和部署的应用程序之间有什么区别吗?这种行为可能来自哪里以及我如何解决这个问题?
据我所知,这不是(直接)编码问题,因为输入正确显示了“Umlaute”......由于调试时无法重现,我很难弄清楚到底出了什么问题。
编辑:虽然输入在日志中看起来不错(即变音符号显示正确),但我意识到它们在控制台中看起来不正确:
<th>Nationalit?ñt:</th>
我目前正在使用 Mirko 建议的 Normalizer,如下所示:(
Normalizer.normalize(input, Form.NFC);
也尝试使用 NFD)。(SpringBoot-) 控制台和 (logback) 日志输出有何不同?
像变音符号这样的变音符号在 unicode 中通常可以用两种不同的方式表示:作为单个代码点字符或作为两个字符的组合。这不是编码的问题,它可能发生在 UTF-8、UTF-16、UTF-32 等中。 Java 的 equals 方法可能不会考虑等于单码点字符的复合字符,即使它们看起来完全一样。尝试查看您正在比较的字符串的二进制表示,这样您应该能够追踪差异。您还可以使用“Character”类的方法来遍历字符串并打印出所有字符的属性。也许这也有助于找出差异。
在任何情况下,如果您java.text.Normalizer在“等号”的“两侧”都使用,将文本规范化为例如 Unicode 规范化形式 C ,这可能会有所帮助。 这样,应该理顺上述差异,并且字符串应该按预期进行比较。
| 归档时间: |
|
| 查看次数: |
801 次 |
| 最近记录: |