Paper: 2604.20666 Authors: Ioannis E. Livieris, Athanasios Koursaris, et al. Categories: cs.CL, cs.AI
Abstract
Effective retrieval-augmented generation across bilingual Greek-English applications requires embedding models capable of capturing both domain-specific semantic relationships and cross-lingual semantic alignment. Existing multilingual embedding models distribute their representational capacity across numerous languages, limiting optimization for Greek with its morphological complexity.
Key Innovation: ORPHEAS
A specialized Greek-English embedding model trained with a knowledge graph-based fine-tuning methodology applied to a diverse multi-domain corpus. This enables language-agnostic semantic representations while preserving Greek-specific morphological and terminological structures.
Why Specialized Models?
- Multilingual models spread capacity: 100+ languages means less capacity per language
- Greek morphological complexity: Greek has rich inflection and compound word formation
- Domain-specific terminology: Technical terms need specialized encoding
Results
Outperforms state-of-the-art multilingual embedding models on:
- Monolingual Greek retrieval benchmarks
- Cross-lingual Greek-English retrieval benchmarks
Domain-specialized fine-tuning on morphologically complex languages does not compromise cross-lingual retrieval capability.
Takeaways
- Language-pair specialized models can outperform general multilingual models
- Knowledge graph-based fine-tuning enables robust cross-lingual alignment
- Morphological complexity is not a barrier to effective cross-lingual retrieval
论文: 2604.20666 作者: Ioannis E. Livieris, Athanasios Koursaris等 分类: cs.CL, cs.AI
摘要
跨双语希腊语-英语应用的有效检索增强生成需要嵌入模型能够捕捉领域特定的语义关系和跨语言语义对齐。现有多语言嵌入模型将表示能力分布在众多语言上,限制了对具有形态复杂性的希腊语的优化。
关键创新:ORPHEAS
一种专门的希腊语-英语嵌入模型,使用基于知识图的微调方法训练,应用于多样的多领域语料库。这使得语言无关的语义表示成为可能,同时保留了希腊语特有的形态和术语结构。
为什么需要专门模型?
- 多语言模型分散能力:100多种语言意味着每种语言的能力减少
- 希腊语形态复杂性:希腊语有丰富的屈折变化和复合词构成
- 领域特定术语:技术术语需要专门的编码
实验结果
在以下方面优于SOTA多语言嵌入模型:
- 单语希腊语检索基准
- 跨语言希腊语-英语检索基准
对形态复杂语言的领域专门微调不会损害跨语言检索能力。
要点总结
- 语言对专门模型可以优于通用多语言模型
- 基于知识图的微调实现鲁棒的跨语言对齐
- 形态复杂性不是有效跨语言检索的障碍