Hevolve AI: Self-Evolving Multimodal AI Agents

Turn your domain expertise into AI agents that keep learning. Hevolve AI lets experts build multimodal AI systems by talking to them and correcting them in real time, with no code to write.

Key Features

Quick Links

© 2024 Hevolve AI Pvt Ltd. All rights reserved.

← All research
Artificial Intelligence
arXiv (cs.CL) · July 16, 2026

Linear representations of grammaticality in neural language models

Jane Li, Najoung Kim

Here is a question that trips up a lot of students: if a model gives a grammatical sentence a higher probability than a broken one, does that mean it knows grammar? Not necessarily, because probability also tracks how common the words are, how plausible the meaning is, and what the model knows about the world. Grammaticality and likelihood get tangled together.

So rather than compare probabilities, the authors look inside the model. Using a technique called mass-mean probing, they check whether grammatical and ungrammatical sentences land in clearly separated regions of the model's internal representation space. They find that separation is robust across many pretrained models, holds up even after accounting for those confounding sentence properties, and generalizes across grammatical phenomena and, to a lesser degree, across languages.

This is drawn from the abstract, which cuts off partway, so read the paper for the methods and caveats.

From the arXiv (cs.CL) abstract

Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that…


More Artificial Intelligence papers