The legal profession necessitates a multidimensional approach that involves synthesizing an in-depth comprehension of a legal issue with insightful commentary based on personal experience, combined with a comprehensive understanding of pertinent legislation, regulation, and case law, in order to deliver an informed legal solution. The present offering with generative AI presents major obstacles in replicating this, as current models struggle to integrate and navigate such a complex interplay of understanding, experience, and fact-checking procedures. It is noteworthy that where generative AI outputs understanding and experience, which reflect the aggregate of various subjective views on similar topics, this often deflects the model's attention from the crucial legal facts, thereby resulting in hallucination. Hence, this paper delves into the feasibility of three independent LLMs, each focused on understanding, experience, and facts, synthesising as one single ensemble model to effectively counteract the current challenges posed by the existing monolithic generative AI models. We introduce an idea of mutli-length tokenisation to protect key information assets like common law judgements, and finally we interrogate the most advanced publicly available models for legal hallucination, with some interesting results.
翻译:法律职业 necessitates 一种多维方法,该方法需要将法律问题的深入理解与基于个人经验的深刻评论相结合,同时综合对相关立法、法规和判例法的全面把握,以提出明智的法律解决方案。目前,生成式人工智能在复制这一过程方面面临重大挑战,因为现有模型难以整合并驾驭如此复杂的理解、经验和事实核查程序的相互作用。值得注意的是,当生成式人工智能输出理解和经验时,这些输出反映了类似主题上各种主观观点的汇总,这往往会转移模型对关键法律事实的关注,从而导致幻觉。因此,本文探讨了三种独立的大语言模型(LLMs)——分别专注于理解、经验和事实——作为单一集成模型进行合成的可行性,以有效应对当前单一生成式人工智能模型带来的挑战。我们引入了多长度分词化(multi-length tokenisation)概念,以保护普通法判决等关键信息资产,最后我们对最先进的公开可用模型进行了法律幻觉方面的审查,并得出了一些有趣的结果。