1 Introduction
In this article, we discuss some exceptions to the Condition on Extraction Domains, such as (1). In (1), wh movement is permitted from an adjunct clause, despite the fact that adjuncts are typically thought to be islands for extraction.
- (1)
- What is the floweri open [proi to attract ____]?
- (Cf. The flower is open to attract passing pollinators.)
Interestingly, this counter-example to the Condition on Extraction Domains is sensitive to interpretation: while the obligatory control interpretation of (1) permits wh movement out of the adjunct clause, a corresponding non-obligatory control variant, as in (2), blocks such movement (Truswell 2011).
- (2)
- *What is the doori open [proarb to listen to_____]?
- (Cf. The door is open to listen to confessions.)
Taking inspiration from Müller 2010 and Landau 2021, we suggest that (1) and other exceptions to the Condition on Extraction Domains are multiple specifier constructions—treating adjuncts as a kind of specifier. Unlike those previous works, however, we offer a theory of how multiple specifier constructions obviate island effects that also takes into account the significance of interactions between different dependencies, such as movement and control, as in (2).
The analysis is built on the idea that (all) syntactic dependencies make use of an operation: Search (Chomsky 2004). Importantly for the present approach, Search is not blind but is guided by the distribution of checked features in the clause. An algorithm for projecting said features therefore makes predictions about which parts of the structure are accessible to Search and which aren’t, thus producing both transparent and opaque domains. As we will see, the particular algorithm that we propose, combined with the above assumptions about Search, makes nuanced predictions about when a specifier or adjunct will be transparent for extraction or control. One of these predictions is that first specifiers of a head are opaque while second specifiers are transparent. We argue that this theory explains both basic Condition on Extraction Domains effects and exceptions to that condition more successfully and in a broader sense than alternative theories.
An outline of the article is as follows. Section 2 presents the empirical evidence that multiple specifier environments obviate the Condition on Extraction Domains, with examples of wh extraction and parasitic gap licensing across control adjunct boundaries, as well as extraction from specifiers (melting; Müller 2010). Section 3 presents the analysis, which treats these effects through the lens of a theory of probing. According to our proposal, probes use the distribution of projected checked features to guide them to goals, with dependency formation being contingent on successfully finding a goal. Section 4 shows how the analysis captures the phenomena in section 2. Section 5 discusses alternative approaches to Condition on Extraction Domains effects and concludes.
2 Variable transparency in multiple specifier constructions
In this section, we first discuss Robert Truswell’s observation that wh movement out of adjuncts tracks the obligatory–non-obligatory control distinction: control adjuncts that are transparent for wh movement show obligatory control, while control adjuncts that are opaque to wh movement show non-obligatory control. Second, we observe that parasitic gaps also track the obligatory–non-obligatory control distinction. Since Nissenbaum 2000’s independently proposed structures for parasitic gap constructions are also multiple specifier configurations, we suggest that the two observations should receive a common analysis. In both cases, the adjunct appears to be transparent for multiple dependencies—binding of an operator or wh extraction plus obligatory control—and in both cases, the adjunct appears in a multiple specifier environment. Thus the same explanation that accounts for correlations between wh movement and control could extend to parasitic gaps and control. Third, we discuss melting, a phenomenon described extensively in Müller 2010, in which scrambling an object across a subject licenses extraction out of the subject, in violation of the Condition on Extraction Domains. We propose, following Müller and others, that scrambling involves a step of object movement through the edge of vP. Since the subject is also generated at the edge of vP, scrambling creates a multiple specifier environment, which we propose is analogous to the environment seen in the first two case studies. Thus, the same explanation that accounts for extraction from control adjuncts can extend to melting examples.
2.1 Control and adjunct (non)-islands
The distinction between obligatory and non-obligatory control is illustrated by the following.
- (3)
- a.
- The floweri is open [proi to attract passing pollinators].
- b.
- The doori is open [proarb to listen to confessions].
The non-agentive, inanimate subjects in (3) may co-refer with the embedded pro, as in (3a), or not, as in (3b). In the latter case, the embedded pro is interpreted as referring to an arbitrary individual/group who might serve as a listener in this context. Insights from Chomsky 1981, Williams 1992, Landau 2013, and Landau 2021 teach us that an inanimate interpretation for pro requires c-command by a controller, while an animate interpretation for pro does not. To reflect this difference, (3a) is called a case of obligatory control and (3b) a case of non-obligatory control.
In McFadden & Sundaresan 2018 it is argued that obligatory and non-obligatory control, despite appearances like in (3), are in complementary distribution. On that view, a better description of (3) would be that (3a) is a case of genuine control while (3b) is what happens when control cannot be established—that is, (3b) is an elsewhere construction. When the predicate of the adjunct clause is strongly biased toward an interpretation where its subject has sentience—as is the case with listen—and when there is no corresponding controller, a preference for the elsewhere construction emerges.
This view of the obligatory–non-obligatory control distinction is further motivated by the observation that obligatory control adjuncts are transparent for wh movement while non-obligatory control adjuncts are opaque (Truswell 2011):
- (4)
- a.
- What is the floweri open [proi to attract _____]?
- b.
- *What is the doori open [proarb to listen to ______]?
The examples in (4) show that infinitival adjuncts can either be fully transparent for wh movement and control or fully opaque. We propose that this variable transparency of control adjuncts results from a structural ambiguity in their attachment sites.1
Following Landau 2021 and references there, we propose that the control adjuncts under discussion are attached within vP. In addition, we follow Landau in assuming that adjuncts that are ambiguous between an obligatory control interpretation and a non-obligatory control interpretation have an ambiguous position within vP. Landau suggests that the two relevant attachment positions for a control adjunct of this sort are above and below the base position of the external argument, as in (5); each choice is compatible with a different control outcome.
- (5)
- Landau’s obligatory versus non-obligatory control
- a.
- Obligatory control: subject is second specifier
- b.
- Non-obligatory control: subject is first specifier
This proposal suggests that the variable transparency of an adjunct clause is tied to its position. An adjunct that is in the right position to participate in obligatory control is also in the right position to be transparent to wh movement. An adjunct that is in the wrong position to participate in obligatory control will conversely be opaque to wh movement. We now discuss other cases where the position of an adjunct is shown to affect its transparency for dependencies into it.
2.2 Control and parasitic gaps
We have discussed how wh movement out of adjuncts correlates with obligatory control into them. We now observe that this correlation with control is not limited to wh movement. Parasitic gaps inside adjuncts show the same correlation. Observe in (6) that parasitic gaps are possible in obligatory control adjuncts but not in non-obligatory control adjuncts.
- (6)
- a.
- [What direction]i was the flowerj opened to what direction
- [Opi proj in order to attract passing pollinators from Op]?
- b.
- *[What sort of person]i was the doorj opened to what sort of person
- [Opi proarb in order to listen to confessions from Op]?
In section 2.1, we speculated that obligatory control adjuncts are transparent as a consequence of the position that they are merged in. Adjuncts that are adjoined in a position that makes them available for obligatory control are transparent both for control and for wh movement; the same adjunct merged in a different position will be opaque for both control and wh movement, forcing a non-obligatory control interpretation of pro. We propose to take the same basic approach to the contrast in (6).
Nissenbaum 2000 develops a theory of parasitic gap licensing that is compatible with such an approach. For independent reasons, Nissenbaum proposes that parasitic gap–containing adjuncts must merge in a particular position: immediately above the subject that controls pro and immediately below a position occupied by the wh element that licenses the parasitic gap. In other words, parasitic gap–containing adjuncts must be second specifiers of vP, as in the following schema.
- (7)
- Nissenbaum’s structure for parasitic gap licensing
According to Nissenbaum’s analysis, if the adjunct merges as a second specifier, it will be accessible to parasitic gap licensing. Since parasitic gaps are only licensed in obligatory control, not in non-obligatory control, as shown in (6), adjuncts in second specifier position must also be allowed to participate in obligatory control. Non-obligatory control adjuncts must not be in this position, since they block parasitic gap licensing. On this view, if the adjunct is in a different position, it will no longer be accessible for parasitic gap licensing or obligatory control, accounting for the correlation of these different dependencies into the adjunct clause.
We have now seen two different phenomena where control appears to correlate with another kind of dependency into an adjunct clause. On the one hand, we saw that obligatory control adjuncts permit wh extraction out of them while non-obligatory control adjuncts don’t. On the other hand, we saw that obligatory control adjuncts also tolerate parasitic gaps while non-obligatory control adjuncts don’t. In both cases, there is precedent from the literature for tracing accessibility to these dependencies to the position of the adjunct: adjuncts in a particular position are accessible to obligatory control and parasitic gaps, suggesting that adjuncts in other positions are opaque to such dependencies.
Unfortunately, though, there is a conflict. Landau proposed that it was the first specifier position that made an adjunct accessible to obligatory control, while Nissenbaum suggested that it was the second specifier position that made an adjunct accessible to parasitic gap licensing. For these two dependencies to correlate, we need to decide which position is the right one for adjuncts that are transparent to obligatory control, wh movement, and parasitic gaps. We resolve this issue in section 2.4, after discussing a third example of multiple specifier constructions with variable transparency effects.
Before moving on, we want to address an analogy that we are drawing between dependencies of different types. We follow Chomsky 1986, Larson 1988, Postal 1998, and Nissenbaum 2000 in assuming that parasitic gap constructions do not involve across-the-board wh movement out of both matrix and adjunct clauses but rather involve binding of an operator that moves adjunct-internally. Nonetheless, we consider it to be significant that parasitic gaps pattern with wh movement with respect to their relationship to obligatory control. We think their similarity in this respect supports a view of operator binding as being subject to the same locality conditions as wh movement. In sections 3 and 4, we propose a theory of locality that applies uniformly to dependencies such as parasitic gap licensing, wh movement, and control.
2.3 Melting—from adjuncts to specifiers
Müller 2010 discusses a class of exceptions to the Condition on Extraction Domains, which Müller calls melting effects. He observes that external arguments in German and Czech are typically opaque to extraction, as expected for specifiers according to the Condition on Extraction Domains. However, he shows that scrambling an object to the left of the external argument has the effect of making the external argument transparent for extraction. In other words, object scrambling obviates the Condition on Extraction Domains for transitive subjects. This is shown in (8) and (9) for German and Czech respectively. Wh extraction out of the subject is only available when the object appears to its left.2
- (8)
- German wh extraction out of subjects
- a.
- *Was1
- what
- haben
- have
- [DP3 t1
- für
- for
- Bücher]
- books.nom
- [DP2
- den
- the
- Fritz]
- Fritz.acc
- beeindruckt?
- impressed
- Intended: ‘What kind of books impressed Fritz?’
- b.
- Was1
- what
- haben
- have
- [DP2
- den
- the
- Fritz]
- Fritz.acc
- [DP3 t1
- für
- for
- Bücher]
- books.nom
- t2
- beeindruckt?
- impressed
- ‘What kind of books impressed Fritz?’
- (9)
- Czech extraction out of subjects
- a.
- *Stará1
- old.nom
- neudeřila
- hit
- [DP3
- žádná
- no.nom
- t1]
- Petra2.
- Petr.acc
- Intended: ‘No old one hit Petr.’
- b.
- (?)Stará1
- old.nom
- neudeřila
- hit
- Petra2
- Petr.acc
- [DP3
- žádná
- no.nom
- t1] t2.
- ‘No old one hit Petr.’
Importantly, Müller cites evidence from Grewendorf 1989 suggesting that the subject of a psych verb like beeindrucken ‘impress’ is a regular external argument in German and not a VP-internal argument. Thus, it must be a specifier, making (8b) a true counter-example to the Condition on Extraction Domains.
What is surprising about (8) and (9) is that the exact same specifier (was für Bücher/stará žádná) can be opaque in (8/9a) but transparent in (8/9b), solely based on the position of the object. The surface position of the object presumably does not affect the specifierhood of the subject, suggesting that island effects have more to do with local context than with the complement–non-complement distinction.
Müller proposes that the contrast is due to the fact that melting examples like (8b) and (9b) involve multiple specifier constructions. A starting assumption is that object movement proceeds successive-cyclically through the edge of vP. Thus, a scrambled object must arrive in the edge of vP at some point in the derivation, in which case the (b) examples differ from the (a) examples with respect to the total number of specifiers vP has. When no scrambling takes place, the external argument is the only phrase to ever occupy the edge of vP, while in scrambling derivations, vP has two specifiers at some point in the derivation.
Müller presents a phase-based theory of the Condition on Extraction Domains, in which phases can only produce escape hatches as long as they are incomplete. The last-merged element in a phase completes the phase and blocks it from producing an escape hatch. As a result, Müller’s theory predicts that only the last-merged specifier of a phase is opaque for extraction. All earlier-merged material is transparent, including specifiers, because they merge early enough for an escape hatch to be produced. In non-scrambling contexts, the subject is the last-merged specifier of vP, while in scrambling contexts, Müller proposes that the object is the last-merged specifier of vP, making the subject transparent. In other words, his theory requires the following configuration of specifiers in vP in order to capture melting effects.
- (10)
- Müller’s vP in melting contexts: only the highest specifier is opaque → the highest specifier must be the scrambled object
Following Moltmann 1990, Grewendorf & Sabel 1999, McGinnis 1999, and Yoshida 2001, among others—based on evidence from quantifier scope and the position of negation and adverbs—spec-vP is not the final landing site for objects scrambled to the left of subjects; a higher position is, like spec-TP. We therefore expect the object to be able to surface in a position that derives the surface word order OS, regardless of the order of specifiers of vP. Since word order, then, is not conclusive to prove the order of specifiers of vP, we need another metric.
It is sometimes argued that the second movement step in German scrambling has Ā properties (Grewendorf 1988, Webelhuth 1992, and Müller & Sternefeld 1994), in which case we might be able to diagnose the order of specifiers in vP with reconstruction tests. The results of such tests suggest ambiguity. Supposing that the first movement step may have mixed properties (as it targets spec-vP), it might be able to affect binding relations, but the second step, having Ā properties, should not; this then should provide some insight into the order of specifiers in spec-vP. As it turns out, a scrambled anaphoric object may be bound by a subject, as in (11a) (with no Condition C effect), consistent with the order of specifiers SO, while a scrambled quantificational object may bind a pronoun in the subject, as in (11b), motivating the opposite.
- (11)
- a.
- …
- dass
- that
- sichi
- refl.acc
- Hansi
- Hans.nom
- nie
- never
- rasiert.
- shaves
- ‘… that Hans never shaves himself.’
- (Yoshida 2001: (68))
- b.
- …
- weil
- because
- jedeni
- everyone.acc
- seini
- his
- Hund
- dog.nom
- gebissen
- bitten
- hat.
- has
- ‘… because everyone has been bitten by his dog.’
- (Moltmann 1990: (130))
This divergence may indicate that specifier ordering is generally ambiguous, just as the position of adjuncts is, with the choice of position affecting extraction and binding possibilities.3
In sum, we have seen three cases in which multiple specifier constructions feed exceptions to the Condition on Extraction Domains. Adjuncts and specifiers that are normally opaque to extraction or parasitic gap licensing may become transparent if they are in a multiple specifier environment. What remains to be decided is which specifier position gets to be exceptionally transparent to such dependencies and what principles underlie this choice. We will argue, contra Müller and Landau but with Nissenbaum, that it is second specifier position that is transparent.
What follows is a discussion of the position of adjuncts that contain parasitic gaps, contrasting Landau and Nissenbaum’s approaches. The latter but not the former is consistent with the proposal that second specifier position is transparent. Readers who are willing to follow us and Nissenbaum may proceed directly to section 3.
2.4 Which specifiers are transparent?
2.4.1 Proposal
We propose the following. For dependencies like obligatory control and wh movement to cross an adjunct clause boundary, the adjunct must merge as the second specifier of vP, and the subject (controller) must merge as the first specifier of vP. We use the term subject rather than external argument on the assumption that subjects always occupy spec-vP at some point in the derivation. When the subject is an external argument, we assume that it externally merges in spec-vP; when the subject is an internal argument (as in a passive or unaccusative), we assume, following Legate 2003 and Sauerland 2003, that it moves through spec-vP en route to spec-TP. Furthermore, we propose that, whether the subject is internally or externally merged in spec-vP, the same ambiguity is available to adjuncts: adjuncts may merge either before or after the subject forms a specifier of vP, leading to two available configurations. If the derivation chooses the option in which the subject is a first specifier, then we predict the adjunct to be transparent for obligatory control and wh movement.
This configuration of specifiers is consistent with Nissenbaum 2000, which likewise argues that (obligatory control) adjuncts with parasitic gaps merge higher than the subject, creating the configuration that we want:
- (12)
- Nissenbaum’s configuration of specifiers with a parasitic gap–containing obligatory control adjunct such as What did Sue throw out without eating?
However, it is the exact opposite of the order of specifiers proposed by Landau 2021, which suggests that obligatory control can only be established if the adjunct merges below the subject; non-obligatory control arises when the adjunct merges above it:
- (13)
- Landau’s obligatory versus non-obligatory control
- a.
- Obligatory control: subject is second specifier
- b.
- Non-obligatory control: subject is first specifier
Here we will explore each analysis in closer detail and see whether the insights from Landau can be made consistent with our proposal.
2.4.2 Nissenbaum on the position of parasitic gap–containing adjuncts
To start, Nissenbaum 2000 offers two main reasons to put parasitic gap–containing adjuncts where they are in (12), above the subject: (1) constituency tests and binding tests show that they are at least as high as spec-vP, and (2) placing them above the subject allows us to treat adjunction as interpreted via predicate modification. Nissenbaum provides examples like (14), which show that the adjunct out-scopes material internal to the verb phrase, such as the verb and internal arguments.
- (14)
- a.
- John [filed the papers and shelved the books] without reading them.
- b.
- We gave himi a book [without talking to Johni’s mother].
- (Nissenbaum 2000: 37–38, (27a, 29c))
Nissenbaum argues that (14) shows us that the adjunct must be at least as high as spec-vP. He then argues that having the adjunct be a second specifier, where the external argument is a first specifier, gives us the right semantics.
Nissenbaum draws an analogy between parasitic gap–containing adjuncts and constructions with operator movement, such as relative clauses. In a relative clause, an operator moves clause-internally, as in (15a), creating a one place predicate, which modifies the relative noun and gets interpreted within the scope of the higher determiner. In a parasitic gap construction, Nissenbaum argues, following Chomsky 1986, Larson 1988, and Postal 1998, that a similar process happens: there is adjunct-internal operator movement, as in (15b), which creates a one place predicate.
- (15)
- Parasitic gaps as derived by operator movement
- a.
- Relative clause: the book [Op [(that) John threw out t]]
- b.
- Parasitic gap adjunct: [Op [without pro reading t]]
When the adjunct merges in spec-vP, its open argument is saturated by the copy of the wh phrase that moves successive-cyclically through spec-vP. For this analogy to hold, the adjunct must be of type ⟨e, t⟩, and its closest c-commanding phrase must be the wh element.
Nissenbaum also draws on the intuition that adjuncts are interpreted via predicate modification. As a result, in order for the above configuration to be interpretable, the sister of the adjunct must have the same type as it: ⟨e, t⟩ (in need of saturation by the wh element). These conditions are both easily met if the subject is internal to the sister of the adjunct clause, as in (16). Wh movement in the main clause through spec-vP leads to lambda abstraction over the verb phrase. As long as the adjunct merges below the wh element but above the subject, predicate modification and subsequent saturation proceed straightforwardly.
- (16)
- Adding semantics to Nissenbaum’s configuration4
As a result, Nissenbaum’s analysis of parasitic gap licensing requires the opposite of Landau’s specifier ordering. Let’s now explore Landau’s reasons for putting specifiers in the order he does, to see whether any common ground can be found. To preview, Landau independently needs to permit specifiers to move freely to create his configuration, in which case his results may not be contingent on having a particular base order of merge.
2.4.3 Why Landau’s specifiers can be flexible
Landau 2021 takes a significantly different approach to the interpretation of control adjuncts than Nissenbaum 2000. While Nissenbaum proposes that they compose via predicate modification (like other modifiers), Landau proposes that the adjuncts we have been looking at compose with the matrix clause via functional application; he assigns adjunct heads a higher type, which first selects its own clause and then the matrix clause.5 On his view, adjuncts are not uniform in type either: obligatory control adjuncts and non-obligatory control adjuncts are headed by elements of different types with different selectional requirements.
- (17)
- Landau’s semantic types for obligatory and non-obligatory control adjuncts
Due to the semantic types Landau assigns, obligatory control adjuncts have to combine with the main clause before the matrix predicate has merged the subject, as in (18a). By contrast, non-obligatory control clauses have to merge after the subject, as in (18b).
- (18)
- Landau’s obligatory versus non-obligatory control
- a.
- Obligatory control: subject is second specifier
- b.
- Non-obligatory control: subject is first specifier
Landau discusses a challenge from clause-initial obligatory control adjuncts, however. In (19), the obligatory control adjunct appears to c-command the surface position of the subject, despite the fact that its semantics should require it to adjoin to the matrix clause below the subject.
- (19)
- [proi standing on the patio], the plantsi obscure the duck pond.
To capture cases like these, Landau proposes that LF movement of the subject applies, taking the subject from its surface position to a higher position, as shown in (20), to feed the semantics. This movement is both unpronounced and insensitive to syntactic rules that might prevent a specifier from moving to a new specifier position of the same projection.
- (20)
- [proi standing on the patio], the plantsi obscure the duck pond.
Given this amendment, it isn’t obvious that Landau’s semantic approach actually restricts the order of specifiers in the narrow syntax. In other words, if our analysis is right in holding that obligatory control adjuncts are generated above the subject (as in Nissenbaum 2000), which is what allows dependencies into the adjunct clause, Landau’s analysis could be easily made compatible with our approach by just assuming that his LF movement applies whenever necessary to get the semantics right.
In sum, given that obligatory control adjuncts are sometimes clearly above the subject and that any theory of control must be able to account for this, we will simply assume that (12) is the baseline configuration for obligatory control: the controller is the first specifier and the pro-containing adjunct is the second specifier. This follows Nissenbaum’s proposed structure and may not affect Landau’s semantics if we permit LF movement.6
The takeaway is that the syntax enforces the specifier order in (12), whenever the adjunct receives an obligatory control interpretation or permits wh extraction or parasitic gaps. Analogously, we expect the subjects in the melting cases to be transparent only when they are a second specifier of vP.
3 Moving towards paths
Thus far, we have observed that multiple specifier environments can obviate the Condition on Extraction Domains. Furthermore, we have proposed that it is second specifiers, rather than first specifiers, that become exceptionally transparent in these environments. To understand why second specifier position might be special in this way, we suggest that the kinds of dependencies under consideration (wh movement, control, parasitic gap licensing) are governed by locality principles that constrain Search. On this view, there must be something special about second specifiers that makes their contents searchable by a higher probe, while first specifiers are opaque. We propose a particular approach to Search and probing that makes this so, which is grounded in a theory of feature projection. On our view, principles of feature projection make the features of second but not first specifiers visible to higher heads.
3.1 Motivation for the approach taken
The remainder of section 3 will describe the analysis in detail, but we first want to clarify why we chose this approach. The present article is a preliminary attempt to make use of existing probing tools to understand an unusual empirical pattern. The challenge with this kind of pattern is that it refers to specifier number in a way that syntax should not be able to do. We know that grammars are not able to “count” and as such should not be able to treat specifiers differently from each other in a way that references only their number. The present approach therefore tries to single out what is unique about the structural context of a second specifier and exploit that in the analysis. As we will lay out in section 3.3, we hypothesize that second specifiers are special due to how the feature projection algorithm labels projections that already contain a specifier. While there may be other ways to approach this locality profile, we think this one has the advantage of making nuanced predictions with existing mechanisms.
3.2 An interlude on Search and paths
This article develops a novel theory of locality rooted in the notion of paths (Kayne 1981, Pesetsky 1982, and McFadden & Sundaresan 2019). Long distance dependencies, on this approach, must be mediated by a sequence of local dependencies:
- (21)
- Path
- For a probe A to enter into a dependency with a licit goal B, there must be a path of local relationships between A and B.
We propose that underlying this notion of paths is a more general notion of economy in dependency formation. Paths, as we define them, allow the grammar enough information to know whether or not a search procedure, on which long distance dependencies are contingent, will succeed or fail.
Chomsky 2004 and subsequent works suggest that probing involves an operation of “Minimal Search.” Minimal, for the purposes applicable here, means that the search procedure will halt once a match has been found, the hope being that the right specification of the search procedure will capture Relativized Minimality effects (Rizzi 1990). A number of algorithms for Minimal Search have been proposed and discussed in the literature; see Branan & Erlewine 2021 for an overview and Atlamaz 2019, Ke 2019, Preminger 2019, Chow 2022, and Krivochen 2022 for more specific proposals. The basic idea is that nodes in a syntactic tree are sequentially “examined” to see if they match what the probe is specified to look for, with the sequential search algorithm only being able to move to sisters or daughters of failed matches.
Why should Search be minimal? One reason, as Chomsky suggests, might be computational efficiency. Searching the tree involves examining a number of nodes to see if they are a match for the probe. We can define a cost in terms of the number of failed examinations that take place prior to success. Examining as few nodes as possible would be desirable, given that the process of examining a node to see if it is a match bears some computational cost.
With this in mind, consider a scenario in which a probe is destined to fail because there is no corresponding goal anywhere in the structure. The probe’s failure cannot be determined, and the derivation cannot proceed to the next step, until every node in the tree is examined. In terms of computational cost, this is the worst case scenario: every node in the tree must be examined, but doing so does not produce any observable change to the structure.
We suggest that the grammar is designed to avoid costly failed searches of this type.7 In the abstract, we suggest that the grammar is endowed with a set of flags that provide (limited) information about the makeup of a constituent. For instance, in the case of a probe specified for a feature [F], the daughters of a node will only be examined by the search procedure if the node itself bears :
- (22)
- The probing configuration
A desirable consequence of this is that it provides a new perspective on island effects: (some) islands, on this approach, would simply be phrases that lack a flag for the relevant sort of feature. In (23), for example, the internal components of YP, which lacks , will not be subject to Search. Probe–goal relationships for [F] will thus be impossible into YP.
- (23)
- An island configuration
Note that this entails a series of local relationships between a probe and its goal: every node in the sister of the probe that dominates the goal must bear a flag for a feature on that goal.
This proposal raises a number of questions, the most pressing of which we hope to answer. In the next sub-section, we suggest that the flags in (0) are checked selectional features—presumably an independently necessary component of the grammar.8 Checked selectional features provide a record of the derivation: the presence of a checked selectional feature on a maximal projection serves as a flag indicating that either the specifier or complement of that phrase is of a particular sort. The chief innovation here will be an algorithm for determining whether or not checked selectional features are able to project past the maximal projection of the head they originate on. Crucially, this decision is local: it creates paths of local relationships between a probe and a licit goal, in the sense of (21). We show that the theory captures the basics of the classic Condition on Extraction Domains: adjuncts and specifiers are, in the basic case, opaque for extraction, while complements are not. We show also that the theory avoids what we term the “escape hatch problem” for phase-based approaches to the Condition on Extraction Domains, a stipulation that requires adjunct islands to both be phases and consistently lack an edge feature.
3.3 Feature checking and feature projection
As discussed in the previous sub-section, we propose that a notion of paths mediates Search. As Search underlies the establishment of long distance dependencies, paths become pre-conditions for long distance dependencies by extension. One of the consequences of a path-based approach to Search is that there are many scenarios in which Search may fail at the outset, before it has examined any nodes. For example, if the sister of a probe does not bear a flag for the relevant feature, the probe won’t bother to search its sister’s daughters at all, given that there is no path of checked selectional features leading from the probe to any of its sister’s descendant nodes in this case.
Our proposal, which we briefly outline here, is based on the following consequence of this approach. For a probe A to establish a dependency with a goal B, A must as a minimal first step be able to identify the relevant flag on its sister. An algorithm for projecting features checked by B from the head that selects B to the sister of A establishes a series of local relationships that links probe and goal. The projection algorithm thus ensures that if a feature has reached A’s sister, it must have been projected at every node between A’s sister and the goal B, thus creating a path between the two, satisfying the condition in (24). We can thus operate with a shorthand definition of path, shown in (25). Here we will discuss the predictions of the approach for dependencies involving selected elements, such as the long distance dependency in (26), where [•B•] is a selectional feature checked by B (see below). We return to the extraction of unselected elements, for example, adjuncts, in section 4.
- (24)
- Accessibility
- A probe A searching for a goal B may only initiate a search for B if there is a path from A to B.
- (25)
- Path (shorthand)
- There is a path from A to B if A’s sister bears a feature checked by B.
- (26)
- A long distance path from A to B
If for some reason A’s sister does not bear a feature checked by B (i.e., because there is no local B), Search fails at the outset, without examining any nodes in the tree. Thus, Search never applies unnecessarily.
We begin by establishing some assumptions about clause construction, before showing how a modified theory of feature projection creates long distance dependencies according to (24) and (25). Adopting the notation of Müller 2010, we represent the features that drive Merge as in (27). A head that selects for a YP, for example, might bear a feature [•Y•], which is checked when that head (or a projection of it) merges with a YP. In other words, a head X with an unchecked [•Y•] feature that merges with a YP produces a projection XP bearing the checked version of that feature, [•Y•], as in (28).
- (27)
- Merge features
- [•Y•] = an instruction to merge with an element bearing [Y]
- (28)
- Selection for YP
As will become important later, we follow Heck & Müller 2007, Müller 2010, Longenbaugh 2019, and Newman 2024 in assuming that these features may drive any kind of Merge, thus not only External Merge but movement (Internal Merge) as well.
The checked features are inactive in the sense that they may no longer drive syntactic operations. In other words, an element that bears [•X•] does not count as an element that bears [X] and thus cannot feed X Merge. Similarly, [•X•] does not count as a selectional feature, which would require checking by an element bearing [X]. It simply indicates that X Merge has taken place. However, contra Chomsky 2000, Adger 2003, and Asudeh & Potts 2004, among others, we suggest that these checked features do not disappear from the derivation. Instead, we propose that they remain present throughout the computation to serve as a pseudo record of selection. On this view, XP provides more information to higher heads than just its own category feature. It also bears checked Merge features, which tell higher heads something about the elements inside XP, for instance that XP contains a YP in the case of (28).
So far, we have seen how checked features may be projected from a head X to its own maximal projection XP, when it merges with elements it selects for. These features are not deleted and are thus visible to whatever subsequently merges with XP. According to the conditions in (24) and (25), for YP to be accessible to anything beyond XP’s sister, [•Y•] must project past XP. Only if it does so can it ever appear on the sister to a higher probe, making YP accessible to that probe.
We propose that feature projection past a maximal projection is conditioned by the local context of that maximal projection. More specifically, maximal projections whose sisters are what we call indivisible feature bundles get to project their checked selectional features to higher nodes, making their contents accessible to later operations, as stated in (29). Indivisible feature bundles are defined in (30)—they are essentially nodes whose features locally come from a single source.
- (29)
- Checked feature projection
- A feature bundle [•F•], [•G•], … on a maximal projection may project iff its sister is an indivisible feature bundle. Features of non-maximal projections may always project.
- (30)
- An indivisible feature bundle is
- a.
- a feature bundle that comes straight from the lexicon—for example, a terminal node (Matushansky 2006)—or
- b.
- a feature bundle that has projected to a node from only one daughter.
The intuition guiding this approach is that language is binary: at every step of the derivation, feature projection should only involve two bundles of features at a time. If one sister already locally projects two feature bundles, the other cannot project at all. If one sister locally projects one or fewer feature bundles, the other can project one as well.
On this view, lexical items are always indivisible feature bundles. Thus, sisters of lexical items (i.e., complements) will always be permitted to project their checked selectional features. Non-terminal nodes, by contrast, might or might not be indivisible feature bundles; that depends on whether their daughters were allowed to project their features. Thus, sisters to non-terminal nodes (specifiers/adjuncts) might or might not be allowed to project their features.
Recalling the XP maximal projection in (28), we can now calculate the predicted effects of context on whether XP gets to project its [•Y•] feature to higher nodes.
If XP is the element merged first with a head Z (otherwise known as Z’s complement), as in (31), the rule in (29) entails that XP can project its [•Y•] feature, because its sister Z is an indivisible feature bundle.
- (31)
- XP projects [•Y•] to a higher node if it is a complement
If XP is the second-merged element in ZP, that is, the first specifier of ZP, as in (32), its sister is not an indivisible feature bundle. The Z′ sister to XP projects from two daughters: the terminal node (which always projects) and its sister (complements get to project, according to (31)). Since Z′ is not an indivisible feature bundle, XP does not get to project [•Y•] in this context, rendering YP inaccessible to operations external to XP.
- (32)
- X P does not project [•Y•] to a higher node if it is a first specifier
The theory thus accounts for basic Condition on Extraction Domains effects: complements permit sub-extraction but first specifiers do not.
The theory makes a surprising prediction for third-merged elements, however, such as the second specifier XP in (33). Here XP’s sister only locally projects from one daughter. Its sister may bear features that originally came from multiple sources, but the notion of indivisibility that we are pursuing only examines a node’s local context: whether it projects from each of its immediate daughters or not. On this view, the sister to XP in (33) only projects from one daughter because first specifiers cannot project, as we saw in (32).
- (33)
- XP projects [•Y•] to a higher node if it is a second specifier
The first node that dominates a first specifier is therefore an indivisible feature bundle according to (30b), and this licenses projection from a second specifier.9
This approach therefore generates an on-again-off-again profile. If some maximal projection is allowed to project, it often creates a context in which the next-merged maximal projection cannot project. If a maximal projection does not project, it often creates a context in which the next-merged maximal projection can project, and so on. Thus, we expect the time of Merge to determine transparency for higher operations, more so than the complement–specifier distinction. We will leverage this context sensitivity to explain the variable opacity of adjuncts and specifiers in different contexts.10
In sum, we propose that the distribution of checked features on nodes creates paths between probes and goals, where paths are a pre-condition for Search. A probe whose sister bears a feature checked by its goal may initiate Search for that goal, so that each successive node is examined for features checked by the goal until the goal is found. If the probe’s sister has no relevant checked features, Search fails before it starts, avoiding unnecessary and costly searches. We proposed that the distribution of checked features is controlled by the rules of feature projection outlined in (29) and (30): maximal projections may project their checked features if their sisters are indivisible feature bundles but not otherwise. Successive projection of checked features creates paths.
It is important to recall that [•Y•] is not equivalent to the YP that checked it. [•Y•] is a feature that was checked/rendered inactive by an element bearing [Y]. By contrast, YP is a phrase that can check some set of features on a probe, including [•Y•]. Thus, a probe whose sister bears [•Y•] has not “found” YP before searching, because [•Y•] can never feed syntactic operations like Y Merge and Y Agree. The probe must still search for the YP that checked the feature in order to satisfy the probe.
At this point, one might wonder why we don’t simply invoke the feature [Y] in path formation, instead of checked selectional features like [•Y•]. If we had a projection algorithm that projected instances of [Y] (i.e., the property of Y that makes it a goal for a [•Y•] probe), then the resulting paths might look like (34), in which [Y] projects to ZP and beyond.
- (34)
- If [Y] projected instead of [•Y•]
The problem with an alternative like this comes from the literature on pied piping. If a higher head seeks [Y], minimality considerations will cause that probe to attract/agree with the highest instance of [Y] in the structure, rather than the phrase [Y] originated on. The result should therefore be a kind of pied piping: ZP, which contains the original YP, should move/agree, but the YP itself should not be available for sub-extraction, due to minimality. Thus, invoking [Y] instead of [•Y•] would produce a feature percolation type approach to pied piping and would not produce a theory that allows the intended goal to move on its own.11 We therefore need some other feature to mediate path formation, one that does not intervene for the dependency being formed. Checked selectional features do just this, without adding to the existing typology of features: they record what kinds of phrases are dominated by a certain node, without making that node an intervener for the goal of the probe.
Lastly, though the presentation here only discussed the case where one checked feature is projected, we assume that feature projection is wholesale in general. What we mean by this is that multiple checked features on a phrase get projected together as a bundle; a maximal projection cannot selectively project some of its features but not others. The following illustrates wholesaleness of projection.
- (35)
- Projection is wholesale
Because projection is wholesale, we expect a maximal projection to be opaque or transparent to higher operations in a very general sense. The projection rules create long distance dependencies in the way illustrated in (35): if XP is selected as the complement of A, then the newly formed AP not only bears [•X•], indicating the presence of an XP inside it, but also inherits any checked features borne by XP itself, allowing whatever selects for AP to search into it for YP and BP as well.12 A transparent maximal projection is transparent for potentially multiple dependencies across itself: it projected every feature it had, so everything inside it that checked a projected feature is visible to higher heads. An opaque maximal projection is similarly opaque for every imaginable dependency: if a maximal projection projects no features past itself, there can be no paths leading into it. We will see that this all or nothing approach captures correlations between different dependencies that cross adjunct boundaries.
In section 4.2, we will clarify certain assumptions that we make about features that drive syntactic operations, assumptions that underpin much of the discussion that follows. But to foreshadow here: we assume that certain elements bear a [D] feature, which marks them as an argument of a clause. This is the same feature that allows certain elements but not others to satisfy the EPP in English; it is also similar to the Case feature of Van Urk & Richards 2015 and the ϕ feature assumed in Van Urk 2015 and Longenbaugh 2019. We use [wh] as a feature for elements that enter into Ā dependencies.
4 Dependencies through paths
We suggest that variable projection from specifiers and control adjuncts accounts for their variable transparency to obligatory control, wh movement, and parasitic gap licensing.
4.1 Outline of the analysis
In section 2 we observed (following Müller, Nissenbaum, Landau, etc.) that violations of the Condition on Extraction Domains often occur in multiple specifier environments. We furthermore concluded with Nissenbaum, contra Landau and Müller, that second specifiers are always exceptionally transparent.
To understand this pattern, we first propose that wh movement, parasitic gap licensing, and obligatory control involve the establishment of a long distance syntactic dependency through Search. To illustrate the proposal with the wh movement–obligatory control correlation, we propose that wh movement arises when interrogative C searches its complement for a phrase bearing a [wh] feature, which is used to check a [•wh•] feature on itself. Obligatory control likewise involves syntactic dependency formation that is contingent on a successful search, with the sister of a potential binder for pro being subject to search for pro (see Ke 2019 for a similar proposal for reflexive binding). Consequently, there must be a path between pro and its controller for obligatory control to arise and a path between a wh phrase and interrogative C for movement to occur.
As proposed in section 3, such path formation is contingent on successful projection of checked selectional features from the sister of the goal to the sister of the probe. The proposed algorithm for feature projection predicts that features from first specifiers do not project to their mothers, making the contents of first specifiers opaque to higher operations. The features of second specifiers, by contrast, do project, creating paths into them. Thus, second specifiers are predicted to be transparent for dependencies like wh movement and control, accounting for the facts in section 2.
Importantly, when an adjunct/specifier projects its features, there are paths into it for every feature that it projects. As a result, when an adjunct is transparent for control, it is also transparent for other dependencies like wh movement. In what follows, we discuss the specific features we have in mind for each of these dependencies and explore some other multiple specifier contexts.
4.2 Which features/paths?
On the theory developed here, the requirement for there to be a path between two elements “linked” through Search should be seen as a way to ensure that Search will succeed. We now explain how Search interacts with movement to account for the correspondent transparency effects. Recall from section 3.3 that both Internal Merge and External Merge are licensed only when they check [•F•]s. Movement—or Internal Merge—requires an invocation of Search on the sister of the probe for some matching feature, followed by merge of the goal at the root of the tree.
Not only must there be a path of checked features between the two elements in question, but the target of Search must have checked the sort of feature that Search is looking for. In other words, a probe with a feature [•X•] must find a path of [•X•] features to its goal, not just any path of features checked by its goal.13
At this point, one might wonder which features actually establish these paths between the matrix subject and pro and between C and the wh element and how those features get checked/projected. For pro, the answer is straightforward. pro is presumably selected as the external argument of the adjunct clause. We can therefore imagine that it checks a [•D•] feature on adjunct v, which gets projected to the highest node in the adjunct clause, as in (36). We have represented pro as the highest specifier of AdjP on the assumption that pro moves to the edge of its clause (Heim & Kratzer 1998).
- (36)
- pro checks [•D•] on v, which projects to AdjP
As long as control is mediated by a search for DPs, if the [•D•] feature projects to the sister of the matrix subject, the matrix subject may find and control pro.14
For wh elements, we propose that their visibility for wh movement is regulated by the distribution of heads bearing [•wh•]. On the assumption that all phase heads have the necessary machinery for hosting successive-cyclic wh movement, we propose (following Longenbaugh 2019 and Newman 2024) that these probes for movement are represented as Merge-inducing features specified to be checked by wh phrases: [•wh•]. If heads like v, C, and possibly D are endowed with such features, then the manner in which [•wh•] gets checked and projected as [•wh•] could be that wh phrases undergo a step of [•D•]-driven movement to the specifier of an intermediate phase head in the clause, before moving to spec-CP.15
For concreteness, consider the case in (37). Here, v bears both a [•D•] and [•wh•] feature. Internal merge to satisfy [•wh•] is not possible: the complement of v does not bear a [•wh•] feature, so it may not be searched for [wh]. The complement does, however, bear a [•D•] feature, so it may be searched for an element bearing [D], in which case the object will be found. Subsequent merge of the object in spec-vP will check both the [•D•] and the [•wh•] on v. Consequently, the [•wh•] will be able to project higher in the tree from this vP, creating a path between the wh phrase in spec-vP and higher elements in the tree.
- (37)
For the sake of having a concrete analysis, we will adopt this approach here, though it may not be the only possible solution. It does, however, have the advantage that it has precedent in the literature from Canac Marquis 1994, Van Urk & Richards 2015, and Longenbaugh 2017. The first two propose that Ā movement chains involving non-subjects always involve a step of A movement within VP.
Canac Marquis suggests that English Ā movement of objects is analogous to tough movement, involving A movement of a null operator to the inner specifier of the phrase that the element undergoing Ā movement is introduced to as an outer specifier. This step of operator movement allows the moved element to be linked to its gap; the position that the moved object is itself initially merged in is presumably motivated by a need to check the [wh] feature of the relevant projection.16
van Urk & Richards propose that wh movement of objects in Dinka (Nilotic; Sudan) is parasitic on a prior step of A movement to a position low in the clause, based on facts about the exceptional absence of an internal argument in an otherwise obligatorily filled pre-verbal slot, in contexts involving Ā movement of an internal argument. The idea, presented using the feature ontology assumed throughout this article, is that movement to this pre-verbal position may in principle check both [•wh•] and [•D•] features. In the absence of an element that bears [wh], this happens solely to satisfy the needs of [•D•] on the relevant head. When an element bears both, checking of the [•wh•] feature on the head that triggers movement may take place, with the path that licenses movement involving a [•D•] feature.
Longenbaugh, like Canac Marquis, discusses a derivation like this in the context of English tough constructions, though the details of the analyses differ. On Longenbaugh’s view, tough movement involves successive-cyclic composite movement through spec-vP, which has both A and Ā properties. On the present view, this “mixed” movement is represented by the multiple checking of two different features on v: [•D•] and [•wh•].
The proposal that wh movement is mediated by DP movement raises questions about wh movement of non-DPs, such as PP arguments and adjuncts:
- (38)
- a.
- To whom did John first speak?
- b.
- On which day did John first speak?
For adjuncts a fairly straightforward analysis would be to propose that they consistently externally merge in spec-vP, at least in cases where they undergo wh movement. For argument PPs the way forward is less obvious. One possibility is that wh PP arguments, like adjuncts, often have the option of initially merging with a functional head like vP, ensuring that the PP argument checks a wh feature (see Newman 2024 for a proposal along these lines). Another possibility is that PP arguments are required to exit the VP for independent reasons (see Stowell 1981 for such a proposal). Assuming this movement is feature-driven, subsequent merge of a PP argument with vP consequently checks v’s [•wh•] feature in cases where the PP bears [wh].
In sum: pro and wh elements must respectively check [•D•] and [•wh•] if they are to be visible for subsequent Search operations. In the case of pro, this is relatively trivial: pro checks a [•D•] when it is initially merged. In the case of wh elements, it means that the wh element must first undergo movement for independent reasons to an intermediate projection, with checking of [•wh•] on this intermediate position taking place as a side effect. Only after movement to such a position will there be a path of [•wh•] features to the wh element, rendering it visible for subsequent searches.
Before moving on, we want to address another dimension of this issue of which features to invoke in path formation. Here we have focused on the type of feature, for example whether it is of category [D] or [wh]. However, one could ask about tokens of features as well. Does the derivation track which DP a given [•D•] corresponds to?
The two possible answers make different predictions when the matrix verb is transitive. In a transitive clause, regardless of the position of the external argument within spec-vP, a [•D•] from the internal argument should project to every node within vP, including the sister of the adjunct. If path formation only cares about the presence of [•D•] on nodes and doesn’t care whether that feature was actually created by the controller, then we might expect all adjuncts that modify transitive clauses to be obligatory control adjuncts. However, non-obligatory control is available in (39), suggesting that there is a derivation available with no path to the adjunct.17
- (39)
- The bus has a seat belt [proarb to wear].
This result could teach us either of two things: maybe (i) we need checked features to be indexed with the element that checked them, with paths being sensitive to these indices, or maybe (ii) heads that select for DPs suppress existing instances of [•D•] from their sisters, breaking the chain of [•D•] until they introduce their own arguments. For concreteness, we tentatively assume the first possibility, but we are optimistic that the proposal could also work without indices.
4.3 Stacked adjuncts: a loose end
It is worth tying up a loose end. As noted by Nissenbaum 2000, one instance of wh movement may license parasitic gaps in more than one adjunct:
- (40)
- Who will you hire ____ [after interviewing ______ ] [if they recommend _____ strongly]?
As pointed out by a reviewer, an analysis like (1) for such stacked adjuncts poses a potential problem for the theory developed here: only one of the two adjuncts on such an analysis would be of the right parity (the property of being an even/odd specifier) to project its features and license the parasitic gap within.
- (41)
We believe there is good reason, then, to favor an analysis like that in (42) for “stacked” adjuncts of the sort discussed here, following Chomsky 2019.
- (42)
In this structure, the two adjuncts merge with each other prior to their merge with vP. Since the “combined” adjunct is an even parity specifier of vP, it will be able to project its features from this position. Furthermore, since the adjuncts that make up the combined adjunct themselves have a specifier—the null operator that gives rise to the parasitic gap—they will be able to project their features up to the “combined” projection.
This theory also accounts for Nissenbaum’s observation that parasitic gap–containing adjuncts must appear closer to the predicate they are construed with than other adjuncts:18
- (43)
- a.
- *Who will you hire ______ [after interviewing someone else] [if they recommend ______ ]?
- b.
- Who will you hire ___________ [after interviewing _______ ] [if they recommend someone else]?
If Nissenbaum is right that parasitic gap–containing adjuncts are of a distinct semantic type, then they should only be able to conjoin with adjuncts of the same type (i.e., gap-containing adjuncts can conjoin with other gap-containing adjuncts).19 Given this, the adjuncts in (43) must form different specifiers of vP. The inner adjunct may be of even parity provided the Ā-moved object lands above it in vP; the outer adjunct, conversely, may not. It would have to be above both the inner adjunct and the subject but below the wh phrase to license a parasitic gap (according to Nissenbaum). From this position, however, it couldn’t project, blocking a parasitic gap. As a result, when there are two adjuncts, one with a gap and one without, the adjunct with the gap must be the inner adjunct.
What we have seen, then, is that multiple kinds of dependencies that Search plausibly underlies—control and Ā dependencies such as wh movement and binding of null operators—are allowed into adjuncts/specifiers only when those adjuncts/specifiers appear in a particular context. Moreover the theory captures the fact that one and the same adjunct/specifier may be transparent or opaque, given that such elements may merge as second specifiers or not. In section 5 we discuss further implications of the theory we’ve developed and compare it to other theories with comparable empirical coverage.
5 Discussion and conclusion
What we have seen so far is a novel theory of locality in which locality domains are determined by their local context. Section 2 examined several exceptions to the Condition on Extraction Domains and showed that they tend to arise in multiple specifier environments. Section 3 developed a theory of feature projection that captures something like the classical Condition on Extraction Domains but that makes fine-grained predictions about when that condition can be obviated. We showed that the theory is able to account for a number of exceptions to the classical Condition on Extraction Domains and furthermore explains a hitherto-unexplained correlation between extraction from an adjunct and the possibility of a non-obligatory control interpretation for that adjunct. Having motivated and developed this theory of locality, we now compare our approach to previous literature and sketch ways forward for future work.
5.1 Other approaches to the Condition on Extraction Domains
Since Cattell 1976, it has been common to treat specifiers and adjuncts as islands for extraction as a matter of definition. The Condition on Extraction Domains states that any non-complement is opaque for extraction:
- (44)
- Condition on Extraction Domains (Huang 1982, Chomsky 1986, Cinque 1990, and Manzini 1992)
- Movement may not cross a barrier XP, unless XP is a complement.
The Condition on Extraction Domains raises several questions. First, many have shown that it is not exceptionless (see, e.g., Stepanov 2007 for discussion). In particular, we have discussed in this article examples of wh extraction out of adjuncts and specifiers, both of which are clear violations of (44). These counter-examples refute the generality of (44) and suggest that we need a more fine-grained metric for islandhood besides the complement–non-complement distinction. Second, existing attempts to derive (44) face a conceptual disadvantage compared to the present theory.
A popular approach to the Condition on Extraction Domains is to treat adjuncts and specifiers as subject to different rules than complements. For example, Uriagereka 1999, Johnson 2003, Sheehan 2013, and Privoznov 2021 suggest that non-complements must spell out when they merge, rendering their contents inaccessible to further operations.
This approach requires some elaboration to theories of spellout, given that complement clauses are also often proposed to spell out at particular points in the derivation. Phases (including complement clauses) are typically assumed to be opaque to operations external to them after their time of spellout. However, unlike adjuncts/specifiers, phasal complements are thought to have an escape hatch. Elements that move to that escape hatch become accessible to later operations, despite the fact that the phase has spelled out. For the contrast between adjuncts/specifiers and complements to be captured, adjuncts/specifiers must therefore lack an escape hatch.
One could imagine several ways to encode the escape hatch property on a phrase such that complements have escape hatches but adjuncts/specifiers do not. For example, we could stipulate that complementation triggers spellout of the complement of the phase head while adjunction/specifier merge triggers spellout of the entire phrase. Complements therefore have a specifier position that has not spelled out, while adjuncts/specifiers do not. Alternatively we could propose that edges of spelled-out phrases are always accessible but that only certain heads have the ability to attract elements to their edge, with adjuncts/specifiers routinely lacking these edge features, in contrast to complements.
Both of these possibilities require us to stipulate a distinction between complements and non-complements, whereas the theory outlined in this article does not. The present theory treats complements as the element merged first with a head, specifiers as the second-merged element, and so on, reducing the number of ad hoc distinctions we need between different phrases. Moreover, the present theory is able to account for variable islandhood of adjuncts and specifiers, without stipulating special properties of those adjuncts and specifiers. Instead we propose that every phrase (complement or non-complement) is subject to the projection algorithm, which yields different results depending on how many feature bundles are present on each node.
5.2 Conclusion
In sum, this article examined a number of exceptions to the Condition on Extraction Domains and offered a novel theory of locality designed to account for these exceptions. On the proposed approach to locality, a specifier or adjunct is rendered transparent or opaque based on the properties of its sister. This allowed us to explain why the same adjunct or specifier might be transparent in one context but opaque in another: aspects of context more subtle than specifierhood/adjuncthood determine opacity.
The proposed theory made use of a modified projection rule, which conditionally allows features to percolate higher than the maximal projection of a head. When the conditions for projection past the maximal level are met, the contents of that phrase become visible to higher probes, allowing dependencies to target them. Crucially, we proposed that probes cannot search for a goal in the absence of such a path of features (created by feature projection). Phrases whose features get suppressed by the projection rule are therefore predicted to be opaque.
Importantly, the projection rule does not reference the complement–non-complement distinction; it instead references the local feature context of a phrase, that is, the features of its sister. As a result, specifiers/adjuncts are not uniformly predicted to be opaque: only those whose local context suppresses feature projection are opaque, accounting for observed exceptions to the Condition on Extraction Domains.
This approach to locality of course raises many questions, which we have not had space to discuss fully here. For example, we have looked at the transparency/opacity of phrases in their base positions but have not yet considered what the approach should look like for phrases derived by movement. To be more specific, consider the melting cases discussed in Müller 2010. Müller shows that object scrambling licenses extraction out of in situ subjects but not those that have raised to a higher position, as diagnosed by the presence of intervening adverbs:
- (45)
- a.
- Was
- what
- haben
- have
- [den
- the
- Fritz]
- Fritz
- denn
- prt
- [_____
- für
- for
- Bücher]
- books
- beeindruckt?
- impressed
- ‘What sort of books have impressed Fritz?’
- b.
- *Was
- what
- haben
- have
- [den
- the
- Fritz]
- Fritz
- [______
- für
- for
- Bücher]
- books
- denn
- prt
- beeindruckt?
- impressed
This kind of freezing effect raises several questions about our proposal, such as: How do copies of phrases interact with the projection rule? And in cases where movement alters the parity of a specifier, can a path of features created by projection from one position find those elements in their new positions? We hope to explore these and related questions in future research.20
While there are certainly further details to develop, the theory as developed so far already makes interesting and nuanced predictions in a traditionally tricky empirical domain and has the advantage of unifying Condition on Extraction Domains effects with other kinds of locality effects analyzed by a search procedure (e.g., intervention effects). As such, we believe it holds promise for a more unified approach to locality, grounded in the nature of the combinatorial system itself rather than the typology of adjunct/specifier phrases.
Supplementary material
A file containing the appendix can be downloaded at https://doi.org/10.16995/star.23413.s1.
Acknowledgments
First and foremost, we thank our colleagues on the team of the research project Locality and the Argument–Adjunct Distinction: Structure Building Versus Structure Enrichment (funded by the United Kingdom’s Arts and Humanities Research Council and by the German Research Foundation) for their support and feedback on this project: Thomas McFadden, Sandhya Sundaresan, Rob Truswell, and Hedde Zeijlstra. We also benefited greatly from the questions and attention of the participants at the 13th Generative Linguistics in the Old World in Asia conference and audiences at the Massachusetts Institute of Technology and Stony Brook University. Finally, we want to thank the anonymous reviewers for their thoughtful comments and questions, which were instrumental in enabling us to test the proposal in new and important ways. All mistakes are our own.
Competing interests
The authors declare that they have no competing interests.
Notes
- Interestingly, such control adjuncts show a weak island effect: they permit extraction of a DP (under the right circumstances) but not, as in (i), an adjunct.
While we do not offer a theory of weak island–hood here, see the appendix (the supplementary material for this article) for some possible views of weak islands on the present theory. ⮭
- (i)
- *How did the flower open [in order to attract pollinators ______]?
- → A: with a particular UV pattern
- Müller observes that this effect is not limited to extraction of DPs but is also found with PPs:
- (i)
- German wh extraction of PPs out of subjects
- a.
- *[PP1
- Über
- about
- wen]
- whom
- hat
- has
- [DP3
- ein
- a
- Buch
- book.nom
- t1] [DP2
- den
- the
- Fritz]
- Fritz.acc
- beeindruckt?
- impressed
- Intended: ‘About whom did a book impress Fritz?’
- b.
- [PP1
- Über
- about
- wen]
- whom
- hat
- has
- [DP2
- den
- the
- Fritz]
- Fritz.acc
- [DP3
- ein
- a
- Buch
- book.nom
- t1] t2
- beeindruckt?
- impressed
- ‘About whom did a book impress Fritz?’
- (ii)
- Czech extraction of PPs out of subjects
- a.
- *[PP1
- O
- about
- starých
- old
- autech]
- cars
- oslovila
- fascinated
- [DP3
- kniha
- book.nom
- t1]
- Petra2.
- Petr.acc
- Intended: ‘A book about old cars fascinated Petr.’
See also Heycock 1991: chap. 3 for examples of pseudo melting. ⮭- b.
- (?)[PP1
- O
- about
- starých
- old
- autech]
- cars
- oslovila
- fascinated
- Petra2
- Petr.acc
- [DP3
- kniha
- book.nom
- t1] t2.
- ‘A book about old cars fascinated Petr.’
- One might wonder why scrambling cannot be terminal in spec-vP. In other words, why couldn’t the object scramble to an inner specifier and just stay there, melting the subject without changing the word order? We propose that a constraint on string-vacuous scrambling must rule this possibility out (Hoji 1985). ⮭
- FA = functional application. PM = predicate modification. ⮭
- It is important to note that Landau does not propose this analysis for all control adjuncts, just those that he describes as alternating between obligatory and non-obligatory control. Adjuncts that can only receive a strict obligatory control interpretation are treated differently, but we don’t discuss those here. ⮭
- At this point, it is still not obvious whether Landau’s assumptions about the semantics of control can be made compatible with Nissenbaum’s when it comes to parasitic gap licensing in control adjuncts. ⮭
- This point is ultimately orthogonal to the question of whether or not probing may fail without leading to a derivational crash (see Preminger 2014 for some discussion); our position is in principle compatible with either view. Failed searches are consistently costly because they require the entire search space of a probe to be exhausted: each node must be examined to see if it matches the needs of the probe, and the examination of each node is that which bears the cost. In a world where probing is allowed to fail, knowing that a particular instance of it will fail allows the derivation to proceed to the next step without incurring the cost associated with Search. In a world where failed probing leads to crash, knowing that a particular instance of it will fail allows the derivation to be thrown out without incurring the cost. ⮭
- This proposal has some surface similarity to approaches to long distance dependencies in Head-Driven Phrase Structure Grammar and Generalized Phrase Structure Grammar (see, e.g., Gazdar 1981). Both on those approaches and on ours, some information about the relationship between heads and their arguments is projected up the clausal spine to a relevant probe. Our theory uses this idea differently, however, because we assume that dependencies may be formed through movement, contra Head-Driven or Generalized Phrase Structure Grammar, where dependencies are always formed through external merger. As a result, the information that is projected is a checked selectional feature rather than something representing an unmet selectional requirement (e.g., the slash in Generalized Phrase Structure Grammar). In that sense, dependencies in our theory crucially rely on a notion of internal merger: without a previously merged instance of some phrase, a path of the relevant checked selectional features could never be created. ⮭
- The present discussion assumes that Merge and feature projection proceed cyclically: each successive specifier extends the clause, and the feature projection algorithm follows the order of Merge. This may not be a necessary feature of the system, however. If multiple specifiers instead tuck in, as in Richards 1997, we could imagine reformulating the approach so that the feature projection algorithm and its consequences for Search inform representational constraints on movement rather than derivational ones. In a world with tucking in, properties of a phrase’s sister could still inform whether that phrase projects, but projection would not proceed according to the order of Merge but rather according to the resulting constituent structure. ⮭
- A question arises: what happens in the case of “simple” phrases, that is, phrases that have themselves failed to select anything? One approach would be to deny the existence of simple phrases of this sort: on this view, every functional item would enter into some sort of selectional dependency with something else, while lexical items would minimally consist of a root and categorizing head (see Marantz 1997 for a proposal along these lines). Another approach would be to say that simple phrases fail to project features, which could potentially have consequences down the line if the phrase that they are a complement of later takes a specifier: the first specifier in this case should be allowed to project in the way a complement normally would. We leave investigation of these possibilities to future research. ⮭
- For arguments against a feature percolation approach to pied piping, see Heck 2008, Cable 2010, and Cable 2012. ⮭
- It is worth noting that the “wholly transparent”/“wholly opaque” nature of certain domains is not inherent to the theory developed here but only holds if the wholesale nature of projection is assumed. We could, of course, imagine more elaborate theories of feature projection that don’t require wholesaleness of the sort assumed here. The consequence of this would be that some domains would be transparent for some dependencies but not others (see Keine 2019 for some discussion of such patterns). We acknowledge this here as a point of interest for future work but do not develop such elaborations here beyond what has here been said. For now, we will proceed with a fairly simple sub-theory of feature projection, to highlight the—to our mind—interesting fact that the “parity” of specifiers/adjuncts determines whether or not they may project, while acknowledging that a more intricate sub-theory of feature projection might make more intricate predictions. ⮭
- Note that the Search-based system developed here includes but is not limited to canonical probe–goal relationships. For the instances of operator binding and control of pro, we could equally well assume that these elements are subject to a well-formedness constraint requiring them to be local to their binder. The idea, then, would be that the evaluation of this condition would be done through Search. Consequently, we would expect these syntactic relationships to display the same locality profile as probe–goal relationships that trigger movement operations. ⮭
- An equivalent alternative is that whichever feature attracts pro to the edge of the adjunct clause is what establishes the path between the matrix subject and pro. If that feature is also [•D•], however, there is no meaningful difference between the two options. If some other feature is responsible for adjunct-internal movement of pro, then some other feature could be responsible for the control path, but we won’t speculate about what that feature could be here. ⮭
- We do not rule out the possibility that other heads (e.g., V) have [•wh•], in which case wh objects could just check [•wh•] upon being selected by V. We pursue the present option to show that the account is also compatible with a more restrictive theory of the distribution of [•wh•], in which only phase heads have access to such features. ⮭
- One might wonder why, if wh phrases may undergo movement to an intermediate position before moving to spec-CP, this intermediate position is not occupied by a moved phrase in non-interrogative contexts. This is a challenge for theories of successive cyclicity, in which non-interrogative phase heads must bear the necessary machinery to host wh movement but only if there is an interrogative C somewhere else in the structure. We assume that many of these cases will be ruled out at the interfaces. Assuming that a wh-moved element needs to be interpreted within the scope of an interrogative element, if there is no such element, then perhaps the wh phrase isn’t licensed, regardless of where it has moved. As for cases with multiple wh phrases, it may be a matter of conditions on pronunciation that force the lower element to appear in its base position, even if it has covertly moved to a higher position. ⮭
- Thanks to an anonymous reviewer for bringing this example to our attention. ⮭
- A reviewer points out that our theory makes a prediction divergent from Nissenbaum involving constructions with three adjuncts. We expect such sentences as the following, where the innermost and outermost adjunct contain parasitic gaps to the exclusion of the middle adjunct, to be acceptable.
This seems in fact to be the case for a number of the speakers we consulted. We leave a fuller account of the speakers who diverge from this judgment as a topic for future work. ⮭
- (i)
- Who will you hire _______ [without interviewing _______] [if John recommends him] [despite criticizing _______]?
- See also Gould 2020 for discussion of parasitic gap data where Nissenbaum’s generalization appears to hold but his semantics do not. ⮭
- A theory of freezing would help us explain, among other things, why adding non-obligatory control adjuncts to sentences in English doesn’t license melting: English subjects always raise to spec-TP and are thus subject to freezing. ⮭
References
Adger, David. 2003. Core syntax. Oxford University Press.
Asudeh, Ash & Potts, Christopher. 2004. Honorific marking: interpreted and interpretable. Unpublished paper. Presented at: Phi Workshop, McGill University, 28–30 August.
Atlamaz, Ümit. 2019. Agreement, case, and nominal licensing. Doctoral thesis. Rutgers University.
Branan, Kenyon & Erlewine, Michael Yoshitaka. 2021. Locality and minimal search. Unpublished paper. Work conducted at: National University of Singapore.
Cable, Seth. 2010. The grammar of Q: Q-particles, wh-movement, and pied-piping. Oxford University Press.
Cable, Seth. 2012. Pied-piping: introducing two recent approaches. Language and Linguistics Compass 6.12.816–832.
Canac Marquis, Rejean. 1994. A/A-bar chain uniformity. Doctoral thesis. University of Massachusetts, Amherst.
Cattell, Ray. 1976. Constraints on movement rules. Language 52.1.18–50.
Chomsky, Noam. 1981. Lectures on government and binding. Foris Publications.
Chomsky, Noam. 1986. Barriers. MIT Press.
Chomsky, Noam. 2000. Minimalist inquiries: the framework. In: Martin, Roger & Michaels, David & Uriagereka, Juan (editors). Step by step: essays in minimalist syntax in honor of Howard Lasnik. MIT Press. 89–155.
Chomsky, Noam. 2004. Beyond explanatory adequacy. In: Belletti, Adriana (editor). Structures and beyond. Oxford University Press. 104–131.
Chomsky, Noam. 2019. The UCLA lectures. Unpublished paper. Presented at: Department of Linguistics, University of California, Los Angeles, 29 April–2 May.
Chow, Keng Ji. 2022. A novel algorithm for minimal search. Snippets 42.3–5.
Cinque, Guglielmo. 1990. Types of Ā-dependencies. MIT Press.
Gazdar, Gerald. 1981. Unbounded dependencies and coordinate structure. Linguistic Inquiry 12.2.155–184.
Gould, Isaac. 2020. Multiple movement dependencies and parasitic gaps. Canadian Journal of Linguistics/Revue Canadienne de Linguistique 65.1.110–121.
Grewendorf, Günther. 1988. Aspekte der deutschen Syntax [Aspects of German syntax]. Gunter Narr Verlag.
Grewendorf, Günther. 1989. Ergativity in German. Foris Publications.
Grewendorf, Günther & Sabel, Joachim. 1999. Scrambling in German and Japanese: adjunction versus multiple specifiers. Natural Language and Linguistic Theory 17.1.1–65.
Heck, Fabian. 2008. On pied-piping: wh-movement and beyond. Mouton de Gruyter.
Heck, Fabian & Müller, Gereon. 2007. Extremely local optimization. In: Bainbridge, Erin & Agbayani, Brian (editors). Proceedings of the thirty-fourth West Coast Conference on Linguistics. Department of Linguistics, California State University, Fresno. 170–182.
Heim, Irene & Kratzer, Angelika. 1998. Semantics in generative grammar. Blackwell Publishing.
Heycock, Caroline. 1991. Layers of predication: the non-lexical syntax of clauses. Doctoral thesis. University of Pennsylvania.
Hoji, Hajime. 1985. Logical form constraints and configurational structures in Japanese. Doctoral thesis. University of Washington.
Huang, Cheng-Teh James. 1982. Logical relations in Chinese and the theory of grammar. Doctoral thesis. Massachusetts Institute of Technology.
Johnson, Kyle. 2003. Towards an etiology of adjunct islands. Nordlyd 31.1.187–215.
Kayne, Richard S. 1981. ECP extensions. Linguistic inquiry 12.1.93–133.
Ke, Hezao. 2019. The syntax, semantics and processing of agreement and binding grammatical illusions. Doctoral thesis. University of Michigan.
Keine, Stefan. 2019. Selective opacity. Linguistic Inquiry 50.1.13–62.
Krivochen, Diego. 2022. The search for Minimal Search: a graph-theoretic view. Unpublished paper. Work conducted at: University of Oxford.
Landau, Idan. 2013. Control in generative grammar: a research companion. Cambridge University Press.
Landau, Idan. 2021. A selectional theory of adjunct control. MIT Press.
Larson, Richard K. 1988. Light predicate raising. Center for Cognitive Science, Massachusetts Institute of Technology.
Legate, Julie Anne. 2003. Some interface properties of the phase. Linguistic Inquiry 34.3.506–516.
Longenbaugh, Nicholas. 2017. Composite a/a’-movement: evidence from English tough-movement. Unpublished paper. Work conducted at: Massachusetts Institute of Technology. http://ling.auf.net/lingbuzz/003604.
Longenbaugh, Nicholas. 2019. On expletives and the agreement-movement correlation. Doctoral thesis. Massachusetts Institute of Technology.
Manzini, Rita. 1992. Locality, a theory and some of its empirical consequences. MIT Press.
Marantz, Alec. 1997. No escape from syntax: don’t try morphological analysis in the privacy of your own lexicon. University of Pennsylvania Working Papers in Linguistics 4.2.201–225.
Matushansky, Ora. 2006. Head movement in linguistic theory. Linguistic Inquiry 37.1.69–109.
McFadden, Thomas & Sundaresan, Sandhya. 2018. Reducing pro and PRO to a single source. Linguistic Review 35.3.463–518.
McFadden, Thomas & Sundaresan, Sandhya. 2019. Deriving selective opacity via path-based locality. Handout. Presented at: syntax mini-course, University of Cambridge, 22 November. https://www.sndrsn.org/copy-of-agree-ment.
McGinnis, Martha. 1999. A-scrambling exists! University of Pennsylvania Working Papers in Linguistics 6.1.20.
Moltmann, Friederike. 1990. Scrambling in German and the specificity effect. Unpublished paper. Work conducted at: Massachusetts Institute of Technology.
Müller, Gereon. 2010. On deriving CED effects from the PIC. Linguistic Inquiry 41.1.35–82.
Müller, Gereon & Sternefeld, Wolfgang. 1994. Scrambling as A-bar movement. In: Corver, Norbert & van Riemsdijk, Henk (editors). Studies on scrambling. Mouton de Gruyter. 331–385.
Newman, Elise. 2024. When arguments merge. MIT Press.
Nissenbaum, Jonathan W. 2000. Investigations of covert phrase movement. Doctoral thesis. Massachusetts Institute of Technology.
Pesetsky, David Michael. 1982. Paths and categories. Doctoral thesis. Massachusetts Institute of Technology.
Postal, Paul Martin. 1998. Three investigations of extraction. MIT Press.
Preminger, Omer. 2014. Agreement and its failures. MIT Press.
Preminger, Omer. 2019. What the PCC tells us about “abstract” agreement, head movement, and locality. Glossa 4.1.13.
Privoznov, Dmitry. 2021. A theory of two strong islands. Doctoral thesis. Massachusetts Institute of Technology.
Richards, Norvin Waldemar. 1997. What moves where when in which languages? Doctoral thesis. Massachusetts Institute of Technology.
Rizzi, Luigi. 1990. Relativized minimality. MIT Press.
Sauerland, Uli. 2003. Intermediate adjunction with A-movement. Linguistic Inquiry 34.2.308–313.
Sheehan, Michelle. 2013. The resuscitation of CED. In: Kan, Seda & Moore-Cantwell, Claire & Staubs, Robert (editors). NELS 40: proceedings of the fortieth annual meeting of the North East Linguistic Society. 2.135–150.
Stepanov, Arthur. 2007. The end of CED? Minimalism and extraction domains. Syntax 10.1.80–126.
Stowell, Timothy Angus. 1981. Origins of phrase structure. Doctoral thesis. Massachusetts Institute of Technology.
Truswell, Robert. 2011. Events, phrases, and questions. Oxford University Press.
Uriagereka, Juan. 1999. Multiple spell-out. In: Epstein, Samuel David & Hornstein, Norbert (editors). Working minimalism. MIT Press. 251–282.
van Urk, Coppe. 2015. A uniform syntax for phrasal movement: a case study of Dinka Bor. Doctoral thesis. Massachusetts Institute of Technology.
van Urk, Coppe & Richards, Norvin. 2015. Two components of long-distance extraction: successive cyclicity in Dinka. Linguistic Inquiry 46.1.113–155.
Webelhuth, Gert. 1992. Principles and parameters of syntactic saturation. Oxford University Press.
Williams, Edwin. 1992. Adjunct control. In: Larson, Richard & Iatridou, Sabine & Lahiri, Utpal & Higginbotham, James (editors). Control and grammar. Springer. 297–322.
Yoshida, Mitsunobu. 2001. Scrambling in German and Japanese from a minimalist point of view. Linguistic Analysis 30.1–2.93–116.