在python pandas中计算两列之间的许多相同的单词

use*_*235 5 python pandas

我们假设我在python pandas中有下表

friend_description  friend_definition
    James is dumb      dumb dude
    Jacob is smart     smart guy
    Jane is pretty     she looks pretty
    Susan is rich      she is rich
Run Code Online (Sandbox Code Playgroud)

在这里,在第一行中,'dumb'一词包含在两列中.在第二行中,'smart'包含在两列中.在第三行中,'pretty'包含在两列中,在最后一行中,'is'和'rich'包含在两列中.我想创建以下列:

friend_description  friend_definition      word_overlap    overlap_count
    James is dumb      dumb dude              dumb             1
    Jacob is smart     smart guy              smart            1
    Jane is pretty     she looks pretty       pretty           1
    Susan is rich      she is rich            is rich          2
Run Code Online (Sandbox Code Playgroud)

我可以使用for循环来手动定义带有这些东西的新列,但我想知道pandas中是否有一个函数可以使这种类型的操作更加平滑.

Max*_*axU 4

处理此类字符串时,简单列表理解似乎是最快的方法:

\n\n
In [112]: df['word_overlap'] = [set(x[0].split()) & set(x[1].split()) for x in df.values]\n\nIn [113]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [114]: df\nOut[114]:\n  friend_description friend_definition word_overlap  overlap_count\n0      James is dumb         dumb dude       {dumb}              1\n1     Jacob is smart         smart guy      {smart}              1\n2     Jane is pretty  she looks pretty     {pretty}              1\n3      Susan is rich       she is rich   {rich, is}              2\n
Run Code Online (Sandbox Code Playgroud)\n\n

单身的apply(..., axis=1):

\n\n
In [85]: df['word_overlap'] = df.apply(lambda r: set(r['friend_description'].split()) &\n    ...:                                          set(r['friend_definition'].split()),\n    ...:                                axis=1)\n    ...:\n\nIn [86]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [87]: df\nOut[87]:\n  friend_description friend_definition word_overlap  overlap_count\n0      James is dumb         dumb dude       {dumb}              1\n1     Jacob is smart         smart guy      {smart}              1\n2     Jane is pretty  she looks pretty     {pretty}              1\n3      Susan is rich       she is rich   {rich, is}              2\n
Run Code Online (Sandbox Code Playgroud)\n\n

apply().apply(..., axis=1)方法:

\n\n
In [23]: df['word_overlap'] = (df.apply(lambda x: x.str.split(expand=False))\n    ...:                         .apply(lambda r: set(r['friend_description']) & set(r['friend_definition']),\n    ...:                                axis=1))\n    ...:\n\nIn [24]: df['overlap_count'] = df['word_overlap'].str.len()\n\nIn [25]: df\nOut[25]:\n  friend_description friend_definition word_overlap  overlap_count\n0      James is dumb         dumb dude       {dumb}              1\n1     Jacob is smart         smart guy      {smart}              1\n2     Jane is pretty  she looks pretty     {pretty}              1\n3      Susan is rich       she is rich   {is, rich}              2\n
Run Code Online (Sandbox Code Playgroud)\n\n

针对 40.000 行 DF 的计时:

\n\n
In [104]: df = pd.concat([df] * 10**4, ignore_index=True)\n\nIn [105]: df.shape\nOut[105]: (40000, 2)\n\nIn [106]: %timeit [set(x[0].split()) & set(x[1].split()) for x in df.values]\n223 ms \xc2\xb1 19.4 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\n\nIn [107]: %timeit df.apply(lambda r: set(r['friend_description'].split()) & set(r['friend_definition'].split()), axis=1)\n3.65 s \xc2\xb1 46.6 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\n\nIn [108]: %timeit df.apply(lambda x: x.str.split(expand=False)).apply(lambda r: set(r['friend_description']) & set(r['friend_definition']),\n     ...: axis=1)\n4.63 s \xc2\xb1 84.7 ms per loop (mean \xc2\xb1 std. dev. of 7 runs, 1 loop each)\n
Run Code Online (Sandbox Code Playgroud)\n