Digital note-taking is gaining popularity, offering a durable, editable, and easily indexable way of storing notes in the vectorized form, known as digital ink. However, a substantial gap remains between this way of note-taking and traditional pen-and-paper note-taking, a practice still favored by a vast majority. Our work, InkSight, aims to bridge the gap by empowering physical note-takers to effortlessly convert their work (offline handwriting) to digital ink (online handwriting), a process we refer to as Derendering. Prior research on the topic has focused on the geometric properties of images, resulting in limited generalization beyond their training domains. Our approach combines reading and writing priors, allowing training a model in the absence of large amounts of paired samples, which are difficult to obtain. To our knowledge, this is the first work that effectively derenders handwritten text in arbitrary photos with diverse visual characteristics and backgrounds. Furthermore, it generalizes beyond its training domain into simple sketches. Our human evaluation reveals that 87% of the samples produced by our model on the challenging HierText dataset are considered as a valid tracing of the input image and 67% look like a pen trajectory traced by a human.
翻译:数字笔记以其耐用、可编辑、易索引的向量化存储(数字墨水)形式日益普及。然而,这种笔记方式与仍受绝大多数人青睐的传统纸笔笔记之间存在显著鸿沟。我们的研究InkSight旨在通过赋能物理笔记用户轻松将作品(离线手写)转换为数字墨水(在线手写)来弥合鸿沟,我们称此过程为“去渲染”。此前相关研究聚焦于图像的几何特性,导致模型在训练领域外的泛化能力有限。我们的方法结合了阅读与书写的先验知识,使得在缺乏大量难以获取的配对样本情况下仍能训练模型。据我们所知,这是首项能有效去渲染任意照片中具有多样化视觉特征与背景的手写文本的工作。此外,该模型还能泛化至训练领域外的简单草图。人工评估表明,在具有挑战性的HierText数据集上,模型生成的样本中有87%被视为输入图像的有效描摹,67%的样本轨迹与人类手写笔迹一致。