A Crucial Parameter for Rank-Frequency Relation in Natural Languages

Abstract

f r-α · (r+γ)-β has been empirically shown more precise than a na\"ive power law f r-α to model the rank-frequency (r-f) relation of words in natural languages. This work shows that the only crucial parameter in the formulation is γ, which depicts the resistance to vocabulary growth on a corpus. A method of parameter estimation by searching an optimal γ is proposed, where a ``zeroth word'' is introduced technically for the calculation. The formulation and parameters are further discussed with several case studies.

0

Discussion (0)

Sign in to join the discussion.

Loading comments…