Scientific press briefings are a valuable information source. They consist of alternating expert speeches, questions from the audience and their answers. Therefore, they can contribute to scientific and fact-based media coverage. Even though press briefings are highly informative, extracting statements relevant to individual journalistic tasks is challenging and time-consuming. To support this task, an automated statement extraction system is proposed. Claims are used as the main feature to identify statements in press briefing transcripts. The statement extraction task is formulated as a four-step procedure. First, the press briefings are split into sentences and passages, then claim sentences are identified through sequence classification. Subsequently, topics are detected, and the sentences are filtered to improve the coherence and assess the length of the statements. The results indicate that claim detection can be used to identify statements in press briefings. While many statements can be extracted automatically with this system, they are not always as coherent as needed to be understood without context and may need further review by knowledgeable persons.
翻译:科学新闻发布会是一种宝贵的信息来源,其内容交替包含专家发言、观众提问及问答环节,因此有助于推动科学及基于事实的媒体报道。尽管新闻发布会信息丰富,提取与记者个人任务相关的陈述仍具有挑战性且耗时。为支持此任务,本文提出了一种自动陈述提取系统。该系统以主张作为核心特征,用于识别新闻发布会实录中的陈述。陈述提取任务被表述为四个步骤:首先,将新闻发布会内容拆分为句子和段落;其次,通过序列分类识别含有主张的句子;随后进行主题检测,并通过筛选句子以改善连贯性并评估陈述长度。结果表明,主张检测可用于识别新闻发布会中的陈述。尽管该系统能自动提取大量陈述,但提取的陈述在脱离语境时未必具备足够的连贯性,仍需领域专家进一步审核。