小编Mik*_*nen的帖子

如何在Java中将JTextPanes/JEditorPanes html内容清理为字符串?

我试图从JTextPane获得漂亮(清理)的文本内容.以下是来自的示例代码JTextPane:

JTextPane textPane = new JTextPane ();
textPane.setContentType ("text/html");
textPane.setText ("This <b>is</b> a <b>test</b>.");
String text = textPane.getText ();
System.out.println (text);
Run Code Online (Sandbox Code Playgroud)

文字看起来像这样JTexPane:

这是一个考验.

我得到这种打印到控制台:

<html>
  <head>

  </head>
  <body>
    This <b>is</b> a <b>test</b>.
  </body>
</html>
Run Code Online (Sandbox Code Playgroud)

我使用过substring()和/或replace()编码,但使用起来很不舒服:

String text = textPane.getText ().replace ("<html> ... <body>\n    , "");
Run Code Online (Sandbox Code Playgroud)

是否有任何简单的函数<b>从字符串中删除除标签(内容)之外的所有其他标签?

有时在内容周围JTextPane添加<p>标签,所以我也想摆脱它们.

像这样:

<html>
  <head>

  </head>
  <body>
    <p style="margin-top: 0">
      hdfhdfgh
    </p>
  </body>
</html>
Run Code Online (Sandbox Code Playgroud)

我想只获得带有标签的文字内容:

This <b>is</b> …
Run Code Online (Sandbox Code Playgroud)

html java string jtextpane

2
推荐指数
1
解决办法
2761
查看次数

在Windows中使用Java读取UTF-8格式的xml -file会出现"IOException:2字节UTF-8序列的无效字节2".-错误

我的Java程序有问题.我如何读取具有"UTF-8"编码的xml -file.程序在Kubuntu中正常工作,但我在Windows中不起作用.两个操作系统都正确编写xml -file,但解析在Windows中出现异常错误.

String XMLFile = "ÄÄKKÖSET.xml"
Document doc = DocumentBuilderFactory.newInstance().newDocumentBuilder().parse(new File (XMLFile));
Run Code Online (Sandbox Code Playgroud)

这是我需要解析的xml -file:

<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<deck created="04/04/2011">
  <title>ääkköset</title>
  <code>ÄÄKKÖSET</code>
  <description>ääkköset</description>
  <author>ääkköset</author>
  <cards nextCardID="1">
    <card color="#1364F9" id="0">
      <question>ÄÄKKÖSET</question>
      <answer>ÄÄKKÖSET</answer>
    </card>
  </cards>
</deck>
Run Code Online (Sandbox Code Playgroud)

如何在Windows中使用Java读取xml -file而不会得到"IOException:2字节UTF-8序列的无效字节2".-错误?

提前致谢!

java xml parsing utf-8

0
推荐指数
1
解决办法
2675
查看次数

标签 统计

java ×2

html ×1

jtextpane ×1

parsing ×1

string ×1

utf-8 ×1

xml ×1