We present the InterviewBot that dynamically integrates conversation history and customized topics into a coherent embedding space to conduct 10 mins hybrid-domain (open and closed) conversations with foreign students applying to U.S. colleges for assessing their academic and cultural readiness. To build a neural-based end-to-end dialogue model, 7,361 audio recordings of human-to-human interviews are automatically transcribed, where 440 are manually corrected for finetuning and evaluation. To overcome the input/output size limit of a transformer-based encoder-decoder model, two new methods are proposed, context attention and topic storing, allowing the model to make relevant and consistent interactions. Our final model is tested both statistically by comparing its responses to the interview data and dynamically by inviting professional interviewers and various students to interact with it in real-time, finding it highly satisfactory in fluency and context awareness.
翻译:我们提出了InterviewBot,该系统能够动态地将对话历史与定制化主题整合到连贯的嵌入空间中,从而与申请美国大学的外国学生进行10分钟的混合领域(开放与封闭式)对话,以评估其学术与文化适应性。为了构建基于神经网络的端到端对话模型,我们自动转录了7,361条人机面试录音,其中440条经过人工校正以进行微调与评估。为克服基于Transformer的编码器-解码器模型在输入/输出长度上的限制,我们提出了两种新方法——上下文注意力机制与主题存储机制,使模型能够实现相关且一致的交互。我们通过两种方式对最终模型进行测试:一是统计性地将其响应与面试数据进行对比,二是通过邀请专业面试官及不同学生与其进行实时交互,结果发现该模型在流畅性与上下文感知方面令人高度满意。