Ney Implementation of Word Based Statistical Language Models
Abstract. In this paper we present an efficient data structure for storing trigram, bigram and unigram counts. The amount of memory required has been reduced by 53 % compared to straightforward approaches. The average access time for retrieving information
IMPLEMENTATION OF WORD BASED STATISTICAL LANGUAGE MODELSFrank Wessel, Stefan Ortmanns and Hermann NeyLehrstuhl fur Informatik VI, RWTH Aachen, University of Technology, D-52056 Aachen, Germany, Phone/Fax:+49 241 f8021616/8888219g, E-mail: wessel@informatik.rwth-aachen.de trigram, bigram and unigram counts. The amount of memory required has been reduced by 53% compared to straightforward approaches. The average access time for retrieving information from the data structure has also slightly been reduced. Based upon this special data structure we have implemented several types of language models and applied them to the North American Business (NAB '94) recognition task. We show that both, the perplexity and the error rate could be reduced compared to the o cial NAB '94 trigram language model.
Abstract. In this paper we present an e cient data structure for storing
1 INTRODUCTIONThe main task of statistical language modelling is to provide a speech recognition system with the a-priori probabilities for a word sequence w1:::wN . In order to be able to compute the widely used bigram and trigram language models, we have to count how often a trigram or bigram, i.e. a word triple or a word pair, has been seen in a training corpus. We can then compute the probability estimate for the trigram u; v; w as: N (u; v; w) p(wju; v )=; N (u; v ) with N (u; v; w) denoting the number of times the trigram u; v; w has occured in the training corpus. To overcome the well-known zero frequency problem, some sort of discounting must be applied to the relative freqencies. The words are usually replaced by word indices which correspond to their position in a lexically sorted vocabulary. Using this text representation, there are several approaches to compute the relative frequency of a trigram:{ The whole training corpus can then easily be stored in a one-dimensional array. Whenever the probability of a speci c trigram is needed, its frequency can simply be calculated by counting the occurrences of the event in the corpus. Assigning two bytes for each word we need 480 MByte to store the whole NAB '94 Corpus consisting of about 240 million words. It is useful to introduce another array, in which the next occurence of a word in the training corpus is stored. The memory cost then rises up to 1.4 GByte. 55
Abstract. In this paper we present an efficient data structure for storing trigram, bigram and unigram counts. The amount of memory required has been reduced by 53 % compared to straightforward approaches. The average access time for retrieving information
corpus size bigrams: ( 1 ( trigrams: ( 1 (
Table 1. Statistics for the NAB'94 Corpus.N u; v< N u; v>
N u; v
N u; v; w
< N u; v; w>
N u; v; w
)=1 ( ) 256 ) 255 )=1 ( ) 256 ) 255<<
240 million words 51 888 422 4 848 006 100 779 41 885 919 17 956 245 71 907
{ The second approach computes the smoothed relative frequencies of all tri-
grams beforehand and stores these probability estimates, herewith avoiding the time consuming computations during the speech recognition process. The trigram probabilities can be easily retrieved using the word indices of the words in the trigram. As the main disadvantage, experiments which require modi cations of the count
s N (u; v; w) (e.g. computing Leaving-One-Outprobabilities on the training corpus) can no longer be performed. Another disadvantage is the still large amount of memory needed, when no cut-o s are applied to the counts. of the trigrams instead of their probabilities. Nevertheless, the memory requirement is very high as presented below. We will show that by using the special structure of the counts the memory cost and the average access time can be reduced.
{ The third approach which we have decided to follow is to store the counts
2 STORING THE COUNTSA rather straigthforward solution to the storing problem is the following: Unigrams, bigrams and trigrams are stored in one array each. Starting with the rst word u in the trigram u; v; w, we look for the second word v in the list of successors which make up a certain part in the bigram list. Applying the same scheme to the trigram list, we can easily nd the word w and the corresponding count N (u; v; w) using binary search. With this approach the memory cost sums up to 420 MByte. Assuming, that the trigram frequencies are almost identical for training and testing, we can easily compute the average number of accesses to the data structure which is needed to nd a speci c trigram: X X N (u; v) ld (SUC(u)+ 1)? 1+ ld (SUC(u; v)+ 1)? 1]; (1) with SUC(u; v) denoting the number of di erent words w succeeding the wordpair u; v in the training and N being the training corpus length. With this approximation we obtain the following results: Finding a trigram takes 17.8 search accesses on average and 29 accesses in the worst case. 56u vN
Abstract. In this paper we present an efficient data structure for storing trigram, bigram and unigram counts. The amount of memory required has been reduced by 53 % compared to straightforward approaches. The average access time for retrieving information
UnigramsBigrams
N(u,v) = 1Trigrams
N(u,v,w) = 11
u
20000N(u,vN(u,v,w)w
Abstract. In this paper we present an efficient data structure for storing trigram, bigram and unigram counts. The amount of memory required has been reduced by 53 % compared to straightforward approaches. The average access time for retrieving information
special case there only exists one successor w following v and the position of the word w in the trigram singleton list e …… 此处隐藏:8984字,全部文档内容请下载后查看。喜欢就下载吧 ……
相关推荐:
- [法律文档]苏教版七年级语文下册第五单元教学设计
- [法律文档]向市委巡视组进点汇报材料
- [法律文档]绵阳市2018年高三物理上学期第二次月考
- [法律文档]浅析如何解决当代中国“新三座大山”的
- [法律文档]延安北过境线大桥工程防洪评价报告 -
- [法律文档]激活生成元素让数学课堂充满生机
- [法律文档]2014年春学期九年级5月教学质量检测语
- [法律文档]放射科标准及各项计1
- [法律文档]2012年广州化学中考试题和答案(原版)
- [法律文档]地球物理勘查规范
- [法律文档]《12系列建筑标准设计图集》目录
- [法律文档]2018年宁波市专技人员继续教育公需课-
- [法律文档]工会委员会工作职责
- [法律文档]2014新版外研社九年级英语上册课文(完
- [法律文档]《阅微草堂笔记》部分篇目赏析
- [法律文档]尔雅军事理论2018课后答案(南开版)
- [法律文档]储竣-13827 黑娃山沟大开挖穿越说明书
- [法律文档]《产品设计》教学大纲及课程简介
- [法律文档]电动吊篮专项施工方案 - 图文
- [法律文档]实木地板和复合地板的比较
- 探析如何提高电力系统中PLC的可靠性
- 用Excel函数快速实现体能测试成绩统计
- 教师招聘考试重点分析:班主任工作常识
- 高三历史选修一《历史上重大改革回眸》
- 2013年中山市部分职位(工种)人力资源视
- 2015年中国水溶性蛋白市场年度调研报告
- 原地踏步走与立定教学设计
- 何家弘法律英语课件_第十二课
- 海信冰箱经销商大会——齐俊强副总经理
- 犯罪心理学讲座
- 初中英语作文病句和错句修改范例
- 虚拟化群集部署计划及操作流程
- 焊接板式塔顶冷凝器设计
- 浅析语文教学中
- 结构力学——6位移法
- 天正建筑CAD制图技巧
- 中华人民共和国财政部令第57号——注册
- 赢在企业文化展厅设计的起跑线上
- 2013版物理一轮精品复习学案:实验6
- 直隶总督署简介




