连接列,仅在存在多个值(非 NA 值)时保留分隔符

问题描述 投票:0回答:1

有谁知道 R 中的一种方法可以连接 n 列,但仅在该行中有值时才保留分隔符?如果您运行下面的示例:

df <- data.frame(
                  name1 = c("Jim","Bob","Sue"),
                  name2 = c("Jane","","Bane"),
                  name3 = c('Conor',"",""),
                  name4 = c("","","Bonor")
                )

df$names <- paste(df$name1,df$name2,df$name3, sep=";")

您将看到分隔符包含在末尾和中间值中,即使单元格为空,输出为:

df =

name1 name2 name3 name4  names
Jim   Jane  Conor        Jim;Jane;Conor;
Bob                      Bob;;;
Sue   Bane        Bonor  Sue;Bane;;Bonor

在单元格为空的情况下,是否有任何方法可以不包含或删除分隔符?达到预期的结果:

df =

name1 name2 name3 name4  names
Jim   Jane  Conor        Jim;Jane;Conor
Bob                      Bob
Sue   Bane        Bonor  Sue;Bane;Bonor
r dataframe concatenation na separator
1个回答
4
投票
library(dplyr)
library(tidyr)

df %>% 
  mutate_all(na_if,"") %>% 
  unite("names", everything(), sep = ";", remove = F, na.rm = T)

#>            names name1 name2 name3 name4
#> 1 Jim;Jane;Conor   Jim  Jane Conor  <NA>
#> 2            Bob   Bob  <NA>  <NA>  <NA>
#> 3 Sue;Bane;Bonor   Sue  Bane  <NA> Bonor

更新:将此解决方案应用于特定列。

我正在修改下面评论中的 akrun 的答案;

df %>% 
  mutate(across(c("name1", "name2", "name3", "name4"), na_if, "", 
                .names = "{.col}_changed")) %>% 
  unite(names, ends_with('_changed'), na.rm = TRUE, sep = ";")

#>   name1 name2 name3 name4          names
#> 1   Jim  Jane Conor       Jim;Jane;Conor
#> 2   Bob                              Bob
#> 3   Sue  Bane       Bonor Sue;Bane;Bonor
© www.soinside.com 2019 - 2024. All rights reserved.