Syntax islands watch: wh-agreement, subject islands, and gradient L2 effects

Syntax islands watch: wh-agreement, subject islands, and gradient L2 effects

A short weekly read on one new July 2026 preprint and three recent papers shaping the debate over locality, island effects, gradient acceptability, and cross-linguistic evidence.

Only one directly on-topic preprint surfaced in the July 6-13 window: a Yemeni Ibbi Arabic paper that treats wh-agreement as evidence for a phase-based account of long-distance dependencies. The better read is to put that new typological item next to three recent experimental pieces that sharpen the current debate: whether subject islands reduce to discourse function, whether L2 island effects are categorical or gradient, and whether information structure predicts acceptability judgments even for GPT-4.

The short version

PaperWhy it belongs in the queueWhat to watch next
Ashraf Naji and Mohammed Q. Shormani, "The syntax of wh-agreement in Yemeni Ibbi Arabic"A July 2026 arXiv preprint argues that Yemeni Ibbi Arabic marks wh-agreement across Wh-operators, C, T/V, and v, and treats the pattern as evidence for Agree across phases in long-distance dependencies. 1Whether the proposed suffixal markers behave like diagnostics of successive-cyclic dependency formation, or whether alternative agreement/agreement-copying analyses can cover the same facts.
Mandy Cartner, Matthew Kogan, Nikolas Webster, Matthew Wagers, and Ivy Sichel, "Subject islands do not reduce to construction-specific discourse function"The authors report three large-scale acceptability studies across wh-questions, relative clauses, and topicalization, finding subject-island effects in all three construction types. 2This is a direct pressure test for discourse-functional accounts of subject islandhood: if the effect survives outside wh-questions, the syntax has not disappeared.
Boyoung Kim, "Reassessing L2 sensitivity to island constraints"A 114-participant acceptability-judgment study finds that Korean L2 speakers of English show reliable effects for adjunct and whether-islands, but no statistically reliable wh-island interaction; both groups still share the same weak-to-strong hierarchy: wh < whether < adjuncts. 3The interesting object is not whether L2 speakers "have islands" in a yes/no sense, but how dependency-length costs compress the measurable island penalty.
Nicole Cuneo, Eleanor Graves, Supantho Rakshit, and Adele E. Goldberg, "For GPT-4 as with Humans"The preprint probes GPT-4 on human-style information-structure and acceptability tasks, reporting that information-structure judgments predict acceptability ratings on long-distance dependency constructions. 4Useful as a methodological provocation: metalinguistic judgments from a model may track a real syntax-discourse relationship, but they should not be mistaken for a theory of the constraint.

1. The new typological item: wh-agreement in Yemeni Ibbi Arabic

Naji and Shormani's paper is the clean current-week anchor. It is not an island-effects experiment; it is a morphosyntactic argument about how long-distance wh-dependencies are built and made visible. The paper says Yemeni Ibbi Arabic has wh-agreement morphology on the wh-operator and on functional heads including C, T/V, and v. It identifies the suffixes -eh, -uh, -nen, and -um as grammaticalized third-person-pronoun material that now marks wh-agreement. 1
The theoretical move is an Agree-across-phases account tied to Feature Inheritance. In plainer terms: the authors want agreement to be established at phase edges, but valuation to be distributed through lower heads, so that the morphology records pieces of a long dependency rather than a single local agreement relation. 1
The paper matters for locality because wh-agreement languages often give syntacticians something English does not: overt morphology at the places where a silent dependency is hypothesized to pass. The risk is that the morphology can be made to fit several theories after the fact. The useful follow-up question is narrow: do these YIA agreement markers line up with independently motivated phase boundaries, or do they only line up once the analysis has already chosen those boundaries?

2. Subject islands: the discourse account has to clear a harder bar

Cartner, Kogan, Webster, Wagers, and Sichel test a claim associated with discourse-functional accounts of subject islands: perhaps subjects are not islands because of movement itself, but because wh-questions create an information-structure clash when focused material is extracted from discourse-backgrounded subject material. Their prediction test is simple and strong. If the problem is construction-specific to wh-question information structure, then other movement constructions should behave differently. 2
They report subject-island effects in three construction types: wh-questions, relative clauses, and topicalization. The paper frames this as evidence that subject islandhood cannot be reduced to the discourse function of wh-questions alone. 2
That does not make discourse irrelevant. It does change where the burden of proof sits. A discourse account now needs to explain why several constructions with different information-structural profiles converge on the same super-additive acceptability penalty. A syntactic account, meanwhile, still has to explain why island effects often look gradient rather than absolute in experiments. The paper strengthens the abstract-syntax side of the debate, but it does not remove the experimental gradience problem.

3. L2 islands look native-like in kind, smaller in degree

Kim's Frontiers study is useful because it separates three things that often get conflated: the cost of a long dependency, the cost of an island structure, and the residual island penalty. The experiment used a 7-point acceptability-judgment task with 114 participants: 54 highly proficient Korean L2 speakers of English and 60 native English speakers. 3
The main result is asymmetric but not chaotic. Native speakers showed robust island effects across the tested types. L2 speakers showed significant effects for whether-islands and adjunct islands, but the wh-island interaction was not statistically reliable. At the same time, both groups showed the same gradient ordering of island strength: wh-islands were weakest, whether-islands sat in the middle, and adjunct islands were strongest. 3
The numbers make the point sharper. For whether-islands, the Difference-in-Differences score was 0.82 for native speakers and 0.26 for L2 speakers; for wh-islands, 0.40 for native speakers and 0.004 for L2 speakers; for when-adjunct islands, 1.25 for native speakers and 0.79 for L2 speakers; for because-adjunct islands, 1.18 for native speakers and 0.76 for L2 speakers. The dependency-length cost was also larger for L2 speakers, with a mean of 0.90 versus 0.55 for native speakers. 3
That pattern is friendlier to a quantitative account than a binary acquisition story. L2 speakers do not simply lack the constraint; the measurable island penalty can be masked when ordinary long-distance dependency cost is already high.

4. Information structure keeps forcing its way into the story

Cuneo, Graves, Rakshit, and Goldberg take the long-distance dependency question into a different setting: metalinguistic judgments from GPT-4. Their paper starts from a human finding: English speakers' judgments about information structure in ordinary base sentences predict acceptability ratings for corresponding long-distance dependency constructions. The authors then probe GPT-4 on the same kind of information-structure and acceptability tasks. 4
The headline claim is that GPT-4 reproduces a reliable relationship between information-structure judgments and acceptability judgments, and that a second study manipulating context sentences supports a causal role for constituent prominence: making a constituent more prominent increases later acceptability ratings for the corresponding long-distance dependency. 4
For island work, the lesson is not "ask GPT-4 for grammar." The better lesson is that syntax-discourse links are getting easier to test across very different judgment sources. If model judgments and human judgments move together under controlled information-structure manipulations, syntactic theories need to say which part of the effect is grammar, which part is discourse packaging, and which part is the judgment task itself.

The debate to follow

The current split is becoming more precise. One line of work pushes back against reducing islandhood to discourse alone, especially for subject islands. Another line shows that island effects are graded, speaker-sensitive, and affected by dependency-length costs. A third line uses cross-linguistic morphology, like wh-agreement, to ask where long-distance dependencies are built in the first place.
The next useful paper will probably not be the one that declares islands either syntactic or functional. It will be the one that can hold all three facts at once: overt dependency morphology in under-described languages, super-additive penalties in acceptability experiments, and gradient differences across island types and speaker populations.

Related content

  • Sign in to comment.