Research Library — Source Index
24 papers downloaded as PDF (validated), plus link-only items. Folder layout mirrors
docs/02-literature-review.mdsections.
arabic-nlp/ (11 files)
| File | Paper | Link |
|---|---|---|
| arabert-2020.pdf | AraBERT (Antoun et al., OSACT/LREC 2020) | https://arxiv.org/abs/2003.00104 |
| arbert-marbert-2021.pdf | ARBERT & MARBERT (Abdul-Mageed et al., ACL 2021) | https://arxiv.org/abs/2101.01785 |
| gate-arabic-embeddings-2025.pdf | GATE Arabic embeddings (Nacar et al., 2025) | https://arxiv.org/abs/2505.24581 |
| allam-sdaia-2024.pdf | ALLaM (SDAIA, 2024) | https://arxiv.org/abs/2407.15390 |
| jais-core42-2023.pdf | Jais (Core42/G42, 2023) | https://arxiv.org/abs/2308.16149 |
| acegpt-2023.pdf | AceGPT (KAUST/CUHKSZ, NAACL 2024) | https://arxiv.org/abs/2309.12053 |
| arabic-llms-survey-2024.pdf | Survey of Arabic LLMs (Mashaabi et al., 2024) | https://arxiv.org/abs/2410.20238 |
| darwish-trec2002-arabic-light-stemming.pdf | Arabic light stemming CLIR (Darwish & Oard, TREC 2002) | https://trec.nist.gov/pubs/trec11/papers/umd.darwish.pdf |
| arabic-dialect-normalization-bert-gpt2-2024.pdf | Dialect→MSA normalization (Alnajjar & Hämäläinen, JDMDH 2024) | https://doi.org/10.46298/jdmdh.13146 |
| arabic-ecommerce-search-challenges-ijlrp-2022.pdf | Arabic e-commerce search challenges (Yadav, IJLRP 2022) | https://www.ijlrp.com/papers/2022/7/1292.pdf |
| sheinfer-arabic-product-category-lrec-osact-2026.pdf | SHEINfer: category inference from Arabic queries (OSACT@LREC 2026) | http://www.lrec-conf.org/proceedings/lrec2026/workshops/osact/pdf/2026.osact-1.10.pdf |
llm-query-expansion/ (5 files)
| File | Paper | Link |
|---|---|---|
| query-expansion-prompting-llms-google-2023.pdf | QE by prompting LLMs (Jagerman et al./Google, 2023) | https://arxiv.org/abs/2305.03653 |
| csqe-eacl2024.pdf | Corpus-Steered Query Expansion (Lei et al., EACL 2024) | https://aclanthology.org/2024.eacl-short.34/ |
| llm-qe-naacl2025.pdf | LLM-QE ranking-preference alignment (Yao et al., NAACL 2025) | https://arxiv.org/abs/2502.17057 |
| best-practices-query-expansion-llms-2024.pdf | Best practices of QE with LLMs (Wang et al., 2024) | https://arxiv.org/abs/2401.06311 |
| knowledge-aware-query-expansion-naacl2025.pdf | Knowledge-aware QE (Xia et al., NAACL 2025) | https://aclanthology.org/2025.naacl-long.216/ |
ecommerce-search/ (6 files)
| File | Paper | Link |
|---|---|---|
| semantic-product-search-amazon-kdd2019.pdf | Semantic Product Search (Nigam et al., Amazon KDD 2019) | https://arxiv.org/abs/1907.00943 |
| esci-shopping-queries-amazon-2022.pdf | Shopping Queries / ESCI benchmark (Reddy et al., 2022) | https://arxiv.org/abs/2206.06588 |
| taobao-longtail-query-rewriting-llm-2024.pdf | LLM long-tail query rewriting, Taobao BEQUE (Peng et al., WWW 2024) | https://arxiv.org/abs/2311.03758 |
| amazon-shopping-intent-query-rewriting-sigir-ecom-2022.pdf | Query rewriting via shopping intent (Zhang et al., SIGIR eCom 2022) | https://sigir-ecom.github.io/ecom22Papers/paper_5298.pdf |
| walmart-sponsored-search-llm-2024.pdf | Sponsored search relevancy with LLM (Rokon et al., Walmart, 2024) | https://sigir-ecom.github.io/eCom24Papers/paper_19.pdf |
| wands-annotation-guidelines.pdf | WANDS annotation guidelines (Wayfair, official PDF from repo) | https://github.com/wayfair/WANDS |
hybrid-retrieval/ (3 files)
| File | Paper | Link |
|---|---|---|
| rrf-cormack-sigir2009.pdf | Reciprocal Rank Fusion (Cormack et al., SIGIR 2009) | https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf |
| weighted-sum-hybrid-bruch-2022.pdf | Hybrid score fusion analysis (Bruch et al., 2022) | https://arxiv.org/abs/2210.11934 |
| splade-formal-2021.pdf | SPLADE sparse expansion (Formal et al., ICTIR 2021) | https://arxiv.org/abs/2107.05720 |
benchmarks/ (3 files)
| File | Paper | Link |
|---|---|---|
| mr-tydi-multilingual-dense-retrieval-2021.pdf | Mr. TyDi multilingual retrieval benchmark incl. Arabic (Zhang et al., MRL 2021) | https://arxiv.org/abs/2108.08787 |
| miracl-multilingual-retrieval-tacl2023.pdf | MIRACL 18-language retrieval dataset incl. Arabic (Zhang et al., TACL 2023) | https://aclanthology.org/2023.tacl-1.63/ |
| mmarco-multilingual-msmarco-2022.pdf | mMARCO: multilingual MS MARCO (Bonifacio et al., 2021) | https://arxiv.org/abs/2108.13897 |
Link-only (paywalled or blocked at download time)
- WANDS (Chen et al., ECIR 2022) — paper page: https://link.springer.com/chapter/10.1007/978-3-030-99736-6_9 · open dataset: https://github.com/wayfair/WANDS
- Improving stemming for Arabic IR (Larkey et al., SIGIR 2002) — https://dl.acm.org/doi/10.1145/564376.564425
- CLE-QR Taobao query rewriting (CIKM 2022) — https://dl.acm.org/doi/10.1145/3511808.3557068
- MAAQR Alipay multi-agent rewriting (SIGIR 2025) — https://dl.acm.org/doi/10.1145/3726302.3731950
- LLM-based QE fails for unfamiliar/ambiguous queries (SIGIR 2025) — https://dl.acm.org/doi/10.1145/3726302.3730222
- Category-Aligned Retrieval for e-commerce (Aliev, arXiv 2025) — https://arxiv.org/abs/2510.21711
Key datasets & models (not PDFs)
- Amazon ESCI data: https://github.com/amazon-science/esci-data
- Wayfair WANDS data: https://github.com/wayfair/WANDS
- Arabic e-commerce search bench: https://huggingface.co/datasets/prestoai/arabic-ecom-search-bench
- GATE model weights: see paper §models; AraBERT/MARBERT on Hugging Face (aubmindlab, UBC-NLP)
- ALLaM-7B GGUF: Hugging Face (Apache-2.0); Jais: https://huggingface.co/core42 ; AceGPT: https://github.com/FreedomIntelligence/AceGPT