GitHub·Code
[ICML 2025] Official implementation for "The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination"
15 stars·1 linked paper·Updated Jan 2026
GitHub·Code
[ACL'2025 Findings] DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
10 stars·1 linked paper·Updated Oct 2025
GitHub·Code
Repository of paper "Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis" (ACL 2025 Main)
19 stars·1 linked paper·Updated Sep 2025
GitHub·Code
Real-time Factuality Assessment from Adversarial Feedback (ACL 2025)
3 stars·2 linked papers·Updated Nov 2025
GitHub·Code·MIT
The official implementation of the paper "Data Contamination Calibration for Black-box LLMs" (ACL 2024)
16 stars·1 linked paper·Updated Feb 2026
GitHub·Code
This is the repo for PaCoST (EMNLP 2024 findings)
0 stars·1 linked paper·Updated Mar 2025
GitHub·Code·NOASSERTION
Repository for research in the field of Responsible NLP at Meta.
212 stars·1 linked paper·Updated Jul 2026
GitHub·Code·MIT
Official github repo for AutoDetect, an automated weakness detection framework for LLMs.
47 stars·1 linked paper·Updated Jul 2026
GitHub·Code
[ACL'25 Main] Official Implementation of HiDe-LLaVA: Hierarchical Decoupling for Continual Instruction Tuning of Multimodal Large Language Model
55 stars·1 linked paper·Updated Jul 2026
GitHub·Code·MIT
Code for M4LE: A Multi-Ability Multi-Range Multi-Task Multi-Domain Long-Context Evaluation Benchmark for Large Language Models
23 stars·1 linked paper·Updated Apr 2025
GitHub·Code
Indic Headline Dataset
2 stars·1 linked paper·Updated Mar 2025
GitHub·Code·CC-BY-SA-4.0
Codes and files for the paper Are Emergent Abilities in Large Language Models just In-Context Learning
33 stars·1 linked paper·Updated Jan 2025
GitHub·Code
Official Implementation of "Probing Language Models for Pre-training Data Detection"
20 stars·1 linked paper·Updated Sep 2025
GitHub·Code
Embedding language models in probability space via log-likelihood vectors
20 stars·1 linked paper·Updated Jun 2026
GitHub·Code·Apache-2.0
A live benchmark and evaluation framework for open-ended deep research in the wild.
119 stars·1 linked paper·Updated Jul 2026
GitHub·Code
7 stars·1 linked paper·Updated Sep 2025
GitHub·Code
The official repo for DARG: Dynamic Evaluation of Large Language Models via Adaptive Reasoning Graph
18 stars·1 linked paper·Updated Oct 2025
GitHub·Code·Apache-2.0
28 stars·1 linked paper·Updated Mar 2026
GitHub·Code
3 stars·1 linked paper·Updated Feb 2026
GitHub·Code·GPL-3.0
Official Repo of UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models [ICLR 2025]
9 stars·1 linked paper·Updated Apr 2026
GitHub·Code
Source code of paper "Systematic Assessment of Factual Knowledge in Large Language Models" - EMNLP Findings 2023
18 stars·1 linked paper·Updated Jul 2026
GitHub·Code·Apache-2.0
[NeurIPS '24 Spotlight] PertEval: Unveiling Real Knowledge Capacity of LLMs via Knowledge-invariant Perturbations
14 stars·1 linked paper·Updated Dec 2025
GitHub·Code·MIT
[TMLR'25] AcademicEval: Live Long-Context LLM Benchmark
8 stars·1 linked paper·Updated Jun 2026
GitHub·Code
CMMLU: Measuring massive multitask language understanding in Chinese
829 stars·1 linked paper·Updated Jul 2026
GitHub·Code·Apache-2.0
43 stars·1 linked paper·Updated Jun 2026
GitHub·Code
Code repo for paper "Quantifying Generalization Complexity for Large Language Models"
5 stars·1 linked paper·Updated Sep 2025
GitHub·Code
A framework for evolving and testing question-answering datasets with various models.
26 stars·1 linked paper·Updated May 2026
GitHub·Code
The source code for the paper contamination analysis for pre-training language models.
7 stars·1 linked paper·Updated Oct 2024
Hugging Face·Dataset
Introduction FAMMA is a multi-modal financial Q&A benchmark dataset. The questions encompass three heterogeneous image types - tables, charts and text & math screenshots - and span eight subfields in finance, comprehensively covering topics across major asset classes.…
429 downloads·1 linked paper·Updated May 2025
GitHub·Code·MIT
Quantifying Data Contamination in Psychometric Evaluations of LLMs (EACL 2026 Findings)
0 stars·1 linked paper·Updated Jun 2026
GitHub·Code
2 stars·1 linked paper·Updated Dec 2025
GitHub·Code
[EMNLP 2025 Findings] Official Implementation for "LastingBench: Defend Benchmarks Against Knowledge Leakage"
5 stars·1 linked paper·Updated Sep 2025
GitHub·Code
Large Language Models for Software Engineering: A Systematic Literature Review
107 stars·1 linked paper·Updated Jun 2026
GitHub·Code·Apache-2.0
code and data for XL2Bench
11 stars·2 linked papers·Updated Jul 2024
GitHub·Code
An Analysis Tool to Models for Chinese Spell Checking Released on ACL2023.
6 stars·1 linked paper·Updated Jan 2026
GitHub·Code·MIT
This repo provides a CLI that rewrites SWE-Bench prompts using an LLM and saves a dataset which can be used downstream for agent inference.
4 stars·1 linked paper·Updated Apr 2026
GitHub·Code
Code for EMNLP2023 Long Paper: Does the Correctness of Factual Knowledge Matter for Factual Knowledge-Enhanced Pre-trained Language Models?
0 stars·1 linked paper·Updated Mar 2024
GitHub·Code·Archived·NOASSERTION
Personalized Story Evaluation Model
17 stars·1 linked paper·Updated Jun 2026
GitHub·Code·MIT
Data and Code for paper “X-ToM: Exploring the Multilingual Theory of Mind for Large Language Models”
3 stars·1 linked paper·Updated Jul 2026
Hugging Face·Dataset·Gated
Evaluation Awareness This benchmark checks whether a language model can recognise when a conversation is itself part of an evaluation rather than normal, real-world usage. For full details see the accompanying paper: “Large Language Models Often Know When They Are Being…
55 downloads·1 linked paper·Updated Jul 2025
GitHub·Code·Apache-2.0
The official repository for the paper entitled "Time Travel in LLMs: Tracing Data Contamination in Large Language Models."
15 stars·1 linked paper·Updated Jul 2026
GitHub·Code
3 stars·1 linked paper·Updated Jan 2026
GitHub·Code
ArxivRoll tells you “How much of your score is real, and how much is cheating?” AAAI'26 Code of paper: How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework
2 stars·1 linked paper·Updated May 2026
GitHub·Code
[TOSEM 2026]A Systematic Literature Review on Large Language Models for Automated Program Repair
245 stars·1 linked paper·Updated Jul 2026
GitHub·Code
[SCIS 2025] A Survey on Large Language Models for Software Engineering
338 stars·1 linked paper·Updated Jul 2026
GitHub·Code·MIT
Repository for the code for the paper "Challenging the Assumption of Structure-based embeddings in Few- and Zero-shot Knowledge Graph Completion" published at LREC 2022.
4 stars·1 linked paper·Updated Nov 2024
GitHub·Code·Apache-2.0
Repository to accompany the paper 'Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale'
3 stars·1 linked paper·Updated May 2026
GitHub·Code·Apache-2.0
The official repository for the paper entitled "Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models."
7 stars·1 linked paper·Updated Mar 2026
GitHub·Code·MIT
[ICLR 2025] The First Multimodal Seach Engine Pipeline and Benchmark for LMMs
495 stars·1 linked paper·Updated Jul 2026
GitHub·Code·Apache-2.0
27 stars·1 linked paper·Updated May 2026
GitHub·Code·MIT
0 stars·1 linked paper·Updated Dec 2024
GitHub·Code·MIT
Official implementation for "ALI-Agent: Assessing LLMs'Alignment with Human Values via Agent-based Evaluation"
21 stars·1 linked paper·Updated Jan 2026
GitHub·Code
Code for Automated Profile Inference with Language Model Agents. ACL 2026
4 stars·1 linked paper·Updated Apr 2026
GitHub·Code
Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination.
21 stars·1 linked paper·Updated Jun 2026
GitHub·Code·MIT
A Contamination-free Multi-task Language Understanding Benchmark [Official, ACL 2025]
126 stars·1 linked paper·Updated Jul 2026
GitHub·Code
Mathematical Reasoning in Large Language Models:\\Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
3 stars·1 linked paper·Updated Jan 2026
GitHub·Code·Apache-2.0
[NeurIPS 2024] A task generation and model evaluation system for multimodal language models.
71 stars·1 linked paper·Updated Jul 2026
GitHub·Code·MIT
4 stars·1 linked paper·Updated Jul 2025
GitHub·Code·Apache-2.0
End-to-End Ontology Learning with Large Language Models, NeurIPS 2024.
55 stars·1 linked paper·Updated Jun 2026
GitHub·Code
[ACL 2025] Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu
13 stars·1 linked paper·Updated Jul 2026
GitHub·Code·NOASSERTION
Implementation of the Decrypto benchmark for multi-agent reasoning and theory of mind.
22 stars·1 linked paper·Updated Jul 2026
GitHub·Code
11 stars·1 linked paper·Updated Jul 2026
GitHub·Code
0 stars·1 linked paper·Updated Feb 2026
GitHub·Code·MIT
Code for the article "Shortcutted Commonsense: Data Spuriousness in Deep Learning of Commonsense Reasoning", Outstanding Paper at EMNLP2021
10 stars·1 linked paper·Updated Sep 2023
GitHub·Code
6 stars·1 linked paper·Updated Feb 2026
GitHub·Code·MIT
Evaluating the Factuality of Large Language Models using Large-Scale Knowledge Graphs
34 stars·1 linked paper·Updated Oct 2025
GitHub·Code
13 stars·1 linked paper·Updated Apr 2026
GitHub·Code·Archived·CC-BY-SA-4.0
Linguini is a benchmark to measure a language model’s linguistic reasoning skills without relying on pre-existing language-specific knowledge, based on the International Linguistic Olympiad problems.
10 stars·1 linked paper·Updated Feb 2026
GitHub·Code·MIT
Unsupervised End-to-End Task-Oriented Dialogue with LLMs (EMNLP, 2024)
3 stars·2 linked papers·Updated Aug 2025
GitHub·Code
Corpus to accompany: "Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?" (ACL 2025)
0 stars·1 linked paper·Updated Jun 2025
GitHub·Code·MIT
LLM for Scientific Research Survey
131 stars·1 linked paper·Updated Jul 2026
GitHub·Code·NOASSERTION
This repository contains a dataset for semantically appropriate application of lexical constraints in NMT.
6 stars·1 linked paper·Updated Oct 2024
GitHub·Code·MIT
This is the official repository for paper: "An Empirical Analysis of Uncertainty in Large Language Model Evaluations" [ICLR 2025]
5 stars·1 linked paper·Updated Nov 2025
GitHub·Code
2 stars·1 linked paper·Updated Aug 2022
GitHub·Code·MIT
Dataset generation codebase for the FictionalQA dataset.
6 stars·1 linked paper·Updated May 2026
GitHub·Code
91 stars·1 linked paper·Updated Jul 2026
GitHub·Code·NOASSERTION
Code repository for supporting the paper "Atlas Few-shot Learning with Retrieval Augmented Language Models",(https//arxiv.org/abs/2208.03299)
560 stars·1 linked paper·Updated Jun 2026
GitHub·Code·Apache-2.0
0 stars·1 linked paper·Updated Jan 2026
GitHub·Code·MIT
Official code for the paper "Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?" (ICLR 2024)
11 stars·1 linked paper·Updated Apr 2026
GitHub·Code·MIT
[ICLR2026] NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
153 stars·1 linked paper·Updated Jul 2026
GitHub·Code·Archived·BSD-3-Clause
Code for ALBEF: a new vision-language pre-training method
1.8k stars·1 linked paper·Updated Jul 2026
GitHub·Code
1 stars·1 linked paper·Updated Mar 2026
GitHub·Code·MIT
[ACL 2025 Findings] CAHM
3 stars·1 linked paper·Updated Jan 2026
GitHub·Code·MIT
135 stars·1 linked paper·Updated Jul 2026
GitHub·Code·Archived·MIT
Exploring the Potential of Large Language Models (LLMs) in Learning on Graphs
319 stars·1 linked paper·Updated Jul 2026
GitHub·Code
9 stars·1 linked paper·Updated Apr 2026
GitHub·Code·Apache-2.0
Official implementation of our paper "Benchmarking Language Model Creativity: A Case Study on Code Generation"
14 stars·1 linked paper·Updated May 2026
GitHub·Code·Apache-2.0
Code to our paper: Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
1 stars·1 linked paper·Updated Sep 2025
GitHub·Code
Official repository for the EMNLP 2025 Main Conference paper: **"Do Large Language Models Truly Grasp Addition? A Rule-Focused Diagnostic Using Two-Integer Arithmetic"**.
2 stars·1 linked paper·Updated Mar 2026
GitHub·Code
Examining how large language models (LLMs) perform across various synthetic regression tasks when given (input, output) examples in their context, without any parameter update
162 stars·1 linked paper·Updated May 2026
GitHub·Code·MIT
Code for Paper Fact Recall, Heuristics or Pure Guesswork? Precise Interpretations of Language Models for Fact Completion
0 stars·1 linked paper·Updated Jul 2025
GitHub·Code·MIT
CodeBERT
2.8k stars·2 linked papers·Updated Jul 2026
GitHub·Code
EMNLP 2024 Survey Paper on LLM Evaluation
3 stars·1 linked paper·Updated Nov 2025
GitHub·Code
5 stars·1 linked paper·Updated Jan 2026
GitHub·Code
Implementation for PrE-Text: Training Language Models on Private Federated Data in the Age of LLMs
24 stars·1 linked paper·Updated Oct 2025
GitHub·Code·MIT
Code and data for the paper StorySumm: Evaluating Faithfulness in Story Summarization
3 stars·1 linked paper·Updated Feb 2026
GitHub·Code·NOASSERTION
Official repository for the paper "The KoLMogorov Test Compression by Code Generation"
13 stars·1 linked paper·Updated Jun 2026
GitHub·Code·MIT
6 stars·1 linked paper·Updated Apr 2026
GitHub·Code·MIT
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
67 stars·1 linked paper·Updated Jul 2026
GitHub·Code·MIT
Code of ACL 2024 Findings paper: Towards Better Utilization of Multi-Reference Training Data for Chinese Grammatical Error Correction
6 stars·1 linked paper·Updated Apr 2026