Research

Publications

2026

  1. AfriqueLLM results comparing base and adapted models across language groups ACL 2026 Oral

    FlagshipMultilingual language models

    Hao Yu Tianyi Xu Michael A. Hedderich Wassim Hamidouche Syed Waqas Zamir David Ifeoluwa Adelani
    ACL 2026 Oral Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
    Abstract

    AfriqueLLM is an open suite of language models adapted to 20 African languages through continued pre-training on 26B tokens. The release includes model weights, data recipes, training configurations, and evaluation commands; systematic experiments across five base models show that task-aligned data composition is the primary driver of downstream gains.

  2. Search-on-Graph reasoning and knowledge-graph navigation workflow KDD 2026
    Jia Ao Sun* Hao Yu* Fabrizio Gotti Fengran Mo Yihong Wu Yuchen Hui
    +3 authors Zhan Su , Lingfeng Xiao , Jian-Yun Nie
    KDD 2026 Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining
    Abstract

    Search-on-Graph follows an observe-think-navigate paradigm in which an LLM selects relations using the available knowledge-graph structure and its complete reasoning history. It outperforms prior methods across six knowledge-graph question-answering benchmarks without task-specific fine-tuning.

2025

  1. TRUTH information-retrieval pipeline highlighting the reranking stage COLM 2025 Workshop
    Hao Yu Shenyang Huang Zachary Yang Maximilian Puelma Touzel Kellin Pelrine Jean-François Godbout
    +1 author Reihaneh Rabbany
    COLM 2025 Workshop COLM 2025 Workshop on Socially Responsible Language Modelling Research
    Abstract

    TRUTH is a domain-adapted reranking approach for misinformation detection that combines supervised fine-tuning and direct preference optimization.

  2. INJONGO multilingual intent and slot-label example in English and isiXhosa ACL 2025 Oral
    Hao Yu Jesujoba Oluwadara Alabi Andiswa Bukula Jian Yun Zhuang En-Shiun Annie Lee Tadesse Kebede Guge
    +15 authors Israel Abebe Azime , Happy Buzaaba , Blessing Kudzaishe Sibanda , Godson Koffi Kalipe , Jonathan Mukiibi , Salomon Kabongo Kabenamualu , Mmasibidi Setaka , Lolwethu Ndolela , Nkiruka Odu , Rooweither Mabuya , Shamsuddeen Hassan Muhammad , Salomey Osei , Sokhar Samb , Dietrich Klakow , David Ifeoluwa Adelani
    ACL 2025 Oral Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
    Abstract

    INJONGO is a multicultural intent-detection and slot-filling benchmark for 16 African languages, built from native-speaker utterances across banking, travel, home, and dining domains.

2024

  1. Dashboard and social-network view from the societal manipulation simulator NeurIPS 2024
    Maximilian Puelma Touzel Sneheel Sarangi Austin Welch Gayatri Krishnakumar Dan Zhao Zachary Yang
    +9 authors Hao Yu , Ethan Kosak-Hine , Tom Gibbs , Andreea Musulan , Camille Thibault , Busra Tugce Gurbuz , Reihaneh Rabbany , Jean-François Godbout , Kellin Pelrine
    NeurIPS 2024
    Abstract

    The rise of AI-driven manipulation poses significant risks to societal trust and democratic processes. Yet, studying these effects in real-world settings at scale is ethically and logistically impractical, highlighting a need for simulation tools that can model these dynamics in controlled settings to enable experimentation with possible defenses. We present a simulation environment designed to address this. We elaborate upon the Concordia framework that simulates offline, ‘real life’ activity by adding online interactions to the simulation through social media with the integration of a Mastodon server. We improve simulation efficiency and information flow, and add a set of measurement tools, particularly longitudinal surveys. We demonstrate the simulator with a tailored example in which we track agents’ political positions and show how partisan manipulation of agents can affect election results.

  2. Illustration of interleaving multilingual training examples EMNLP 2024
    Senyu Li Hao Yu Jessica Ojo David Ifeoluwa Adelani
    EMNLP 2024 Proceedings of the Fourth Workshop on Multilingual Representation Learning (MRL 2024)
    Abstract

    We present our systems for the three tasks and five languages included in the MRL 2024 Shared Task on Multilingual Multi-task Information Retrieval: (1) Named Entity Recognition, (2) Free-form Question Answering, and (3) Multiple-choice Question Answering. For each task, we explored the impact of selecting different multilingual language models for fine-tuning across various target languages, and implemented an ensemble system that generates final outputs based on predictions from multiple fine-tuned models. All models are large language models fine-tuned on task-specific data. Our experimental results show that a more balanced dataset would yield better results. However, when training data for certain languages are scarce, fine-tuning on a large amount of English data supplemented by a small amount of “triggering data” in the target language can produce decent results.

  3. Auepora evaluation targets for retrieval-augmented generation CCF BigData

    FlagshipSearch and retrieval

    Hao Yu Aoran Gan Kai Zhang Shiwei Tong Qi Liu Zhaofeng Liu
    CCF BigData Big Data: 12th CCF Conference, BigData 2024
    Abstract

    We introduce Auepora, a unified process for evaluating retrieval and generation in RAG systems. The survey compares metrics, datasets, and benchmarks across relevance, accuracy, and faithfulness.

  4. Heatmap showing F1 gains from web retrieval by evidence category COLM 2024
    Jacob-Junqi Tian Hao Yu Yury Orlovskiy Mauricio Rivera Zachary Yang Jean-François Godbout
    +2 authors Reihaneh Rabbany , Kellin Pelrine
    COLM 2024
    Abstract

    This paper develops an agent-based automated fact-checking approach for detecting misinformation. We demonstrate that combining a powerful LLM agent, which does not have access to the internet for searches, with an online web search agent yields better results than when each tool is used independently. Our approach is robust across multiple models, outperforming alternatives and increasing the macro F1 of misinformation detection by as much as 20 percent compared to LLMs without search. We also conduct extensive analyses on the sources our system leverages and their biases, decisions in the construction of the system like the search tool and the knowledge base, the type of evidence needed and its impact on the results, and other parts of the overall process. By combining strong performance with in-depth understanding, we hope to provide building blocks for future search-enabled misinformation mitigation systems.

  5. Persian and English indicators used to construct the ideology dataset LREC EURALI
    Sahar Omidi Shayegan Isar Nejadgholi Kellin Pelrine Hao Yu Sacha Levy Zachary Yang
    +2 authors Jean-François Godbout , Reihaneh Rabbany
    LREC EURALI 2nd Workshop on Resources and Technologies for Indigenous, Endangered and Lesser-resourced Languages in Eurasia @ LREC-COLING 2024
    Abstract

    Large Language Models (LLMs) are now capable of successfully identifying the political beliefs of English-speaking social media users from their posts. However, assessing how LLMs perform in non-English languages remains difficult. In this work, we contribute to this area of research by determining the extent to which LLMs can predict the political ideologies of users on Persian social media. We begin by discussing the challenges associated with defining political parties within the Persian context and propose a solution based on a technique designed for the detection of hyper-partisan ideologies on social media. We create a new benchmark and show the potential and limitations of both open-source and commercial LLMs in classifying the hyper-partisan ideologies of users. We compare these models with smaller fine-tuned ones, both on the Persian language (ParsBERT) and translated data (RoBERTa), and confirm that they considerably outperform generative LLMs in this task. We further demonstrate that the performance of the generative LLMs degrades when classifying users based on their tweets instead of their bios, even if tweets are added as additional information; whereas the smaller fine-tuned models are more robust and achieve similar performance for all input settings. This study represents a first step toward political ideology detection in Persian social media, with implications for future research to understand the dynamics of political conflicts in Iran.

2023

  1. Training-loss curve from supervised fine-tuning experiments arXiv
    Hao Yu* Zachary Yang* Kellin Pelrine Jean Francois Godbout Reihaneh Rabbany
    arXiv arXiv preprint arXiv:2308.10092
    Abstract

    Recent advancements in large language models have demonstrated remarkable capabilities across various NLP tasks. But many questions remain, including whether open-source models match closed ones, why these models excel or struggle with certain tasks, and what types of practical procedures can improve performance. We address these questions in the context of classification by evaluating three classes of models using eight datasets across three distinct tasks: named entity recognition, political party prediction, and misinformation detection. While larger LLMs often lead to improved performance, open-source models can rival their closed-source counterparts by fine-tuning. Moreover, supervised smaller models, like RoBERTa, can achieve similar or even greater performance in many datasets compared to generative LLMs. On the other hand, closed models maintain an advantage in hard tasks that demand the most generalizability. This study underscores the importance of model selection based on task requirements.

  2. SWEET weak-supervision architecture for person-name extraction EMNLP 2023
    Javin Liu* Hao Yu* Vidya Sujaya* Pratheeksha Nair Kellin Pelrine Reihaneh Rabbany
    EMNLP 2023 Findings of the Association for Computational Linguistics: EMNLP 2023
    Abstract

    In this work, we propose a weak supervision pipeline SWEET: Supervise Weakly for Entity Extraction to fight Trafficking for extracting person names from noisy escort advertisements. Our method combines the simplicity of rule-matching (through antirules, i.e., negated rules) and the generalizability of large language models fine-tuned on benchmark, domain-specific and synthetic datasets, treating them as weak labels. One of the major challenges in this domain is limited labeled data. SWEET addresses this by obtaining multiple weak labels through labeling functions and effectively aggregating them. SWEET outperforms the previous supervised SOTA method for this task by 9% F1 score on domain data and better generalizes to common benchmark datasets. Furthermore, we also release HTGEN, a synthetically generated dataset of escort advertisements (built using ChatGPT) to facilitate further research within the community.

  3. Hybrid quantum-classical neural-network and TensorCircuit system diagrams Quantum
    Shi-Xin Zhang Jonathan Allcock Zhou-Quan Wan Shuo Liu Jiace Sun Hao Yu
    +10 authors Xing-Han Yang , Jiezhong Qiu , Zhaofeng Ye , Yu-Qin Chen , Chee-Kong Lee , Yi-Cong Zheng , Shao-Kai Jian , Hong Yao , Chang-Yu Hsieh , Shengyu Zhang
    Quantum Quantum
    Abstract

    TensorCircuit is an open-source tensor-network quantum simulation framework designed for efficient, differentiable NISQ workflows across modern machine-learning backends.