We design alignment-free techniques for comparing a sequence or word, called a target, against a set of words, called a reference. A target-specific factor of a target $T$ against a reference $R$ is a factor $w$ of a word in $T$ which is not a factor of a word of $R$ and such that any proper factor of $w$ is a factor of a word of $R$. We first address the computation of the set of target-specific factors of a target $T$ against a reference $R$, where $T$ and $R$ are finite sets of sequences. The result is the construction of an automaton accepting the set of all considered target-specific factors. The construction algorithm runs in linear time according to the size of $T\cup R$. The second result consists of the design of an algorithm to compute all the occurrences in a single sequence $T$ of its target-specific factors against a reference $R$. The algorithm runs in real-time on the target sequence, independently of the number of occurrences of target-specific factors.
翻译:我们设计无比对技术,用于将目标序列或单词与参考序列集进行比较。目标$T$相对于参考集$R$的特定因子是$T$中单词的因子$w$,该因子不是$R$中任意单词的因子,且$w$的任何真因子都是$R$中某个单词的因子。我们首先解决计算目标$T$相对于参考集$R$的特定因子集合问题,其中$T$和$R$是有限序列集。该结果构建了一个接受所有考虑的目标特定因子的自动机。该构造算法根据$T\cup R$的大小在线性时间内运行。第二个结果为设计算法,用于计算单个序列$T$中相对于参考集$R$的目标特定因子的所有出现位置。该算法在目标序列上实时运行,与目标特定因子出现次数的数量无关。