Edge Rewrite
// HTMLRewriter · presentation

This page was redesigned at the edge.

Cloudflare fetched the original article and streamed it through HTMLRewriter to apply an entirely new visual system without rebuilding the source page.

Jump to content

Draft:Assamese & English Digital Lexicon

From Wikipedia, the free encyclopedia

Assamese & English Digital Lexicon
Official logo of Assamese & English Digital Lexicon
Native name
অসমীয়া আৰু ইংৰাজী ডিজিটেল অভিধান
Available inAssamese, English
FounderNaruttam Boruah
CommercialNo
Launched2026
Content license
Open Access

The Assamese & English Digital Lexicon (অসমীয়া আৰু ইংৰাজী ডিজিটেল অভিধান) is an open-access bilingual online dictionary and lexical database for the Assamese language and English language.[1] Founded by software technologist and researcher Naruttam Boruah and published by the Assam-based software studio Dubori, the platform indexes more than 361,000 lexical entries, consisting of approximately 101,000 native Assamese headwords and 260,000 English translation equivalents.[2]

The project aims to improve digital access and computational language resources for Assamese, an Indo-Aryan language historically classified as a low-resource language in natural language processing. In addition to vocabulary definitions, the platform maintains a specialized corpus of over 813 traditional Assamese proverbs (fokora-jojona) and idiomatic expressions.[3]

Historical background

[edit]

Modern Assamese lexicography began during the British colonial period with the publication of Miles Bronson's A Dictionary in Assamese and English in 1867.[4] This was followed by Hemchandra Barua's etymological dictionary Hemkosh, published in 1900, which established the modern orthographic standard of the language, and later by the Asam Sahitya Sabha's Chandrakanta Abhidhan in 1932.[5] In 2006, early web-based efforts such as Xobdo.org introduced community-contributed online glossaries.[6]

The Assamese & English Digital Lexicon was established in 2026 to consolidate these earlier lexicographical traditions into a modern digital repository with standardized Unicode representation, phonetic transliteration support, and programmatic access for automated language models.[1]

Lexicographical structure

[edit]

The database is compiled using a dual-anchor lexicographical method. Headwords and definitions are cross-referenced with classical 20th-century print corpora—principally the orthographic standards of Hemkosh and Chandrakanta Abhidhan—alongside contemporary digital and print literature to resolve discrepancies in orthography and complex Bengali-Assamese conjunct glyphs.[1]

Entries are classified into a four-tier frequency hierarchy calibrated against general linguistic frequency corpora, including the Oxford 3000 and the Google Web Trillion Word Corpus. This ranking prioritizes everyday conversational and academic vocabulary in search results over archaic or specialized scientific nomenclature. The interface accommodates both native Eastern Nagari script queries and Latin phonetic transliteration.

Infrastructure

[edit]

The public web platform is hosted on serverless edge architecture via Cloudflare Workers, using in-memory distributed key-value storage to provide low-latency lookups globally.[7] In addition to the web platform, an offline mobile edition on Android packages an embedded SQLite database containing the entire dictionary, permitting dictionary queries without an internet connection. Public access to the lexicon is also provided through unauthenticated REST API endpoints and machine-readable manifests following the llms.txt standard.

Datasets and research

[edit]

A curated sub-collection of over 813 Assamese proverbs (fokora-jojona) and idiomatic expressions (jatuwa-thas) has been released as an open benchmark dataset to aid research in computational linguistics and natural language processing for low-resource languages. Each entry provides the original proverb in Assamese script, Romanized transliteration, literal translation, contextual cultural explanation, and contemporary usage examples.

The dataset is deposited in the Zenodo research archive under a permanent DOI (doi:10.5281/zenodo.23032165)[3] and is also mirrored on the Hugging Face machine learning repository.[8]

See also

[edit]

References

[edit]
  1. 1 2 3 "About Assamese & English Digital Lexicon". Dubori. Retrieved 29 September 2026.
  2. ↑ "Assamese-English Digital Lexicon Technical Documentation". Assamese & English Digital Lexicon. Retrieved 29 September 2026.
  3. 1 2 Boruah, Naruttam (2026). "Dubori Assamese Idioms, Proverbs, and Figurative Lexicon Benchmark Dataset". Zenodo. doi:10.5281/zenodo.23032165. Retrieved 29 September 2026.
  4. ↑ Bronson, Miles (1867). A Dictionary in Assamese and English. Sibsagar: American Baptist Mission Press.
  5. ↑ Barua, Hemchandra (1900). Gordon, P. R. (ed.). Hemkosh. Jorhat: Hemkosh Printers.
  6. ↑ Staff Reporter (15 September 2010). "Xobdo online dictionary completes 5 yrs". The Assam Tribune. Retrieved 29 September 2026.
  7. ↑ "Developer API Specification". Assamese & English Digital Lexicon. Retrieved 29 September 2026.
  8. ↑ "Assamese Idioms and Proverbs Dataset". Hugging Face. Dubori / Naruttam Boruah. Retrieved 29 September 2026.
[edit]