Mus*_*h87 2 java similarity sparql jena
我想知道在Java代码中编写此SPARQL查询的简单方法:
select ?input
?string
(strlen(?match)/strlen(?string) as ?percent)
where {
values ?string { "London" "Londn" "London Fog" "Lando" "Land Ho!"
"concatenate" "catnap" "hat" "cat" "chat" "chart" "port" "part" }
values (?input ?pattern ?replacement) {
("cat" "^x[^cat]*([c]?)[^at]*([a]?)[^t]*([t]?).*$" "$1$2$3")
("Londn" "^x[^Londn]*([L]?)[^ondn]*([o]?)[^ndn]*([n]?)[^dn]*([d]?)[^n]*([n]?).*$" "$1$2$3$4$5")
}
bind( replace( concat('x',?string), ?pattern, ?replacement) as ?match )
}
order by ?pattern desc(?percent)
Run Code Online (Sandbox Code Playgroud)
此代码包含在讨论中使用iSPARQL使用相似性度量来比较值.此代码的目的是查找与DBPedia上的给定单词类似的资源.这个方法考虑到我事先知道字符串及其长度.我想知道如何在参数化方法中编写此查询,无论单词和长度如何,它都会返回给我相似性度量.
更新: ARQ - 编写属性函数现在是标准Jena文档的一部分.
看起来您喜欢使用SPARQL的语法扩展来执行查询中更复杂的部分.例如:
SELECT ?input ?string ?percent WHERE
{
VALUES ?string { "London" "Londn" "London Fog" "Lando" "Land Ho!"
"concatenate" "catnap" "hat" "cat" "chat" "chart" "port" "part" }
VALUES ?input { "cat" "londn" }
?input <urn:ex:fn#matches> (?string ?percent) .
}
ORDER BY DESC(?percent)
Run Code Online (Sandbox Code Playgroud)
在这个例子中,假设它<urn:ex:fn#matches>是一个属性函数,它将自动执行匹配操作并计算相似性.
Jena文档很好地解释了如何编写自定义过滤器函数,但是(截至2014年8月7日)几乎没有解释如何实现自定义属性函数.
我将假设您可以将您的答案转换为Java代码以计算字符串相似性,并专注于可以容纳您的代码的属性函数的实现.
实现属性函数
每个属性函数都与特定的属性相关联Context.这允许您将函数的可用性限制为全局或与特定数据集关联.
假设您有一个实现PropertyFunctionFactory(稍后显示),您可以按如下方式注册该函数:
注册
final PropertyFunctionRegistry reg = PropertyFunctionRegistry.chooseRegistry(ARQ.getContext());
reg.put("urn:ex:fn#matches", new MatchesPropertyFunctionFactory);
PropertyFunctionRegistry.set(ARQ.getContext(), reg);
Run Code Online (Sandbox Code Playgroud)
全局和特定于数据集的注册之间的唯一区别在于Context对象来自:
final Dataset ds = DatasetFactory.createMem();
final PropertyFunctionRegistry reg = PropertyFunctionRegistry.chooseRegistry(ds.getContext());
reg.put("urn:ex:fn#matches", new MatchesPropertyFunctionFactory);
PropertyFunctionRegistry.set(ds.getContext(), reg);
Run Code Online (Sandbox Code Playgroud)
MatchesPropertyFunctionFactory
public class MatchesPropertyFunctionFactory implements PropertyFunctionFactory {
@Override
public PropertyFunction create(final String uri)
{
return new PFuncSimpleAndList()
{
@Override
public QueryIterator execEvaluated(final Binding parent, final Node subject, final Node predicate, final PropFuncArg object, final ExecutionContext execCxt)
{
/* TODO insert your stuff to perform testing. Note that you'll need
* to validate that things like subject/predicate/etc are bound
*/
final boolean nonzeroPercentMatch = true; // XXX example-specific kludge
final Double percent = 0.75; // XXX example-specific kludge
if( nonzeroPercentMatch ) {
final Binding binding =
BindingFactory.binding(parent,
Var.alloc(object.getArg(1)),
NodeFactory.createLiteral(percent.toString(), XSDDatatype.XSDdecimal));
return QueryIterSingleton.create(binding, execCtx);
}
else {
return QueryIterNullIterator.create(execCtx);
}
}
};
}
}
Run Code Online (Sandbox Code Playgroud)
因为我们创建的属性函数将列表作为参数,所以我们将其PFuncSimpleAndList用作抽象实现.除此之外,在这些属性函数中发生的大多数魔法都是创建Bindings,QueryIterators并执行输入参数的验证.
验证/结算说明
这应该足以让你继续编写自己的属性函数,如果那是你想要存放字符串匹配逻辑的地方.
未显示的是输入验证.在这个答案中,我假设subject第一个列表参数(object.getArg(0))是bound(Node.isConcrete()),第二个列表参数(object.getArg(1))不是(Node.isVariable()).如果不以这种方式调用您的方法,事情就会爆炸.强化方法(将许多if-else块与条件检查放在一起)或支持替代用例(即:查找值object.getArg(0)是否为变量)留给读者(因为在实现过程中演示,易于测试和显而易见是繁琐的).