← All problems

2024 · open round · D · 20 points

Generally Speaking

NACLO 2024 Problem Committee · syntax · morphology · formal and computational

Think about these two examples:

  1. The hummingbird interviewed the penguin.
  2. The the interviewed penguin hummingbird.

It's unlikely that you have seen either of these examples before. Nonetheless, you can probably tell that example A is a possible English sentence (it's certainly strange, but it is grammatical—that is, it obeys all the rules of English sentence structure), while example B could not be an English sentence.

How are we able to make judgments about sentences we have never seen before? This question relates to the broader topic of generalization—applying what you have learned to new situations or examples. Generalization is an important concept in computational linguistics because we want to build computers that can effectively generalize to sentences they have never seen before.

At NACLO Labs, to study generalization, we have created four different computers that are able to learn from data. The computers are named C1, C2, C3, and C4, and they have never encountered any English sentences before our experiment. We first show the computers the 17 sentences below, which we will refer to as the training set—the set of examples that the computers learn English from. We tell the computers that all 17 of these sentences are grammatical.

Training set:

  1. The linguist met the programmer.
  2. The cartoonist saw the spy.
  3. The watchmaker is asleep.
  4. The happy concierge met the ballerina.
  5. The spy saw the woodcarver.
  6. The spy is knowledgeable.
  7. The main detective visited the ballerina.
  8. The main calligrapher saw the astronaut.
  9. The programmer visited the blacksmith.
  10. The detective met the yodeler.
  11. The spy is happy.
  12. The ballerina is famous.
  13. The linguist is tall.
  14. The talented watchmaker visited the astronaut.
  15. The famous programmer visited the blacksmith.
  16. The cheerful concierge met the ballerina.
  17. The haberdasher saw the linguist.

Note that all of the nouns in these sentences are types of occupations (for example, a haberdasher is someone who sells men's clothing), but for this problem you do not need to know the meanings of these words.

We now show the computers 21 more sentences, which are listed in the table below. For each sentence, we ask each computer whether it thinks the sentence is grammatical (G) or ungrammatical (U). Their answers are recorded in the table. We have also included the answers that a human provided when shown the same sentences.

D1. Some of the cells in the table are empty (marked (a)–(n)). Provide the missing entries.

SentenceC1C2C3C4Human
18.The main detective visited the ballerina.GGGGG
19.The linguist is tall.GGGGG
20.The cartoonist saw the spy.GGGGG
21.The linguist visited the spy.UGGGG
22.The main concierge saw the blacksmith.UGGGG
23.The ballerina is asleep.UGGGG
24.The cartoonist met the programmer.UGGGG
25.The woodcarver visited the programmer.UUGGG
26.The linguist saw the haberdasher.UUGGG
27.The famous concierge visited the watchmaker.UUGGG
28.The linguist is cheerful.UUGGG
29.The calligrapher watchmaker visited met.UUUGU
30.The knowledgeable cheerful talented.UUUGU
31.The tall detective saw the ballerina.UGGG
32.The linguist saw the yodeler.UGGG
33.The visited met linguist programmer.UUGU
34.The yodeler met the woodcarver.UGGG
35.The watchmaker is famous.UGGG
36.The detective met the yodeler.GG
37.The spy is talented.GG
38.The programmer met the cartoonist.GG

D2. The table below contains some more sentences, along with C2 and C3's judgments for those sentences. In each sentence, one word has been hidden from you by replacing it with the string HIDDEN_WORD_1, HIDDEN_WORD_2, or HIDDEN_WORD_3. Write down what each HIDDEN_WORD stands for, choosing from the following options: asleep, happy, or main. Additional notes:

  • You should use each option (asleep, happy, or main) exactly once.
  • Each HIDDEN_WORD stands for the same word in both of the sentences where it appears.
SentenceC2C3
39.The watchmaker is HIDDEN_WORD_1.GG
40.The HIDDEN_WORD_1 concierge visited the astronaut.GG
41.The watchmaker is HIDDEN_WORD_2.UG
42.The HIDDEN_WORD_2 concierge visited the astronaut.GG
43.The watchmaker is HIDDEN_WORD_3.GG
44.The HIDDEN_WORD_3 concierge visited the astronaut.UG

D3. For each sentence in the table above, think about what your own labels would be. Out of these 6 sentences, how many are there where C2 produced the same label as you? How many are there where C3 produced the same label as you?

NOTE: Different people may have different judgments about these sentences. Therefore, as long as you write something for D3, you will get credit—it doesn't matter what you write! We just want you to think a bit about which computer matches your own judgments the best.

Ideally, we would want our computers to make exactly the same predictions as a human would. If you look back at sentences 18 to 38, you will see that C1 and C2 often undergeneralize—there are sentences, such as sentence 25, that these computers label as ungrammatical (U) even though the human called them grammatical (G).

In the other direction, C4 sometimes overgeneralizes—it labels some sentences as grammatical (G) even though the human labeled them as ungrammatical (U). Depending on how you answered question D3, there may also be some examples of undergeneralization or overgeneralization in sentences 39 to 44 as well.

The examples of overgeneralization that we've seen so far were about the domain of sentence structure. Overgeneralization can also happen on the level of individual words. For example, based on pairs of words like walk and walked, or laugh and laughed, a computer might overgeneralize and produce the word thinked (based on think) or bringed (based on bring).

How might word-level overgeneralization affect our interactions with a computer? As an illustration, below are some interactions between a human and a hypothetical robot. The robot has been trained on a very large amount of English text, and it often overgeneralizes when forming words. Each word that resulted from overgeneralization has been hidden from you by being encoded in a cipher, where each letter has been replaced with a different letter. The same code was used for all examples—for example, if X stands for w in the first encoded word, then it will also stand for w everywhere else.

D4. Decode each encoded word from the conversations below. We are not providing the robot's training set, so it will not be possible to understand the robot as precisely as C1, C2, C3, and C4 from the previous part of the problem. However, if you carefully analyze what the robot says and also figure out the code, it should be possible to determine what each encoded word stands for.

Human: How severe will the storm be? Do you think it will interrupt my WiFi?

Robot: The storm will have a high level of severity. An interruption to your WiFi is likely, especially if lightning strikes somewhere nearby.

Human: Are you able to make this picture smaller for me?

Robot: I would be happy to QGTRHUU it for you! The RKGKHENDKEW of the smaller version should make it easier to move around. And I would be happy to enlarge it again later.

Human: That's a strange way to phrase it…but thank you! Also, something looks a little odd about the picture—can you tell what it is?

Robot: The picture is currently UQYETKFQ-DKAPE. I would be happy to apply a mirror filter to fix that. The website where you got the picture has many shortcomings—one of them is that it often distorts images.

Human: [showing a picture] What is the largest object in this picture?

Robot: [pointing at object in picture] That!

Human: Where is the water bottle in the picture?

Robot: [pointing at corner of picture] There!

Human: Who is the tallest person in the picture?

Robot: [pointing at person in picture] EPX!

Human: Can you attempt to schedule an appointment with my hairdresser?

Robot: Yes, I will make this HEEQRMEKXG soon. I will request an appointment TXRQCPQG next week.

Human: [returning after a two-week vacation] It's great to be back in the office! Did anything happen while I was away?

Robot: Not much happened for the first week. But then, WQTEQDCQQB, we got two new clients! They thought our company had many EHUUJXRKGAT, so they were happy to hire us.

An important challenge in computational linguistics is figuring out how to get computers to generalize in the right way. Generalizing can also be a challenge for humans acquiring languages. For example, when acquiring English, children often overgeneralize by producing words such as thinked or bringed.

The rest of the questions are below, under Questions.

TimerOptional. Set whatever limit suits you.

Questions

0 of 27 correct
D2

Using C2's and C3's judgments in sentences 39–44, decide which of asleep, happy, main each HIDDEN_WORD stands for (each option is used exactly once).

HIDDEN_WORD_1 =
HIDDEN_WORD_2 =
HIDDEN_WORD_3 =
D3

Out of sentences 39–44, how many did each computer label the same way you would? (Any answer receives credit — the graders accept whatever you write.)

Number of the 6 sentences where C2 gave the same label as you (0–6):
Number of the 6 sentences where C3 gave the same label as you (0–6):
D4

Decode each of the robot's encoded words. Write the plain English word the robot actually produced.

QGTRHUU =
RKGKHENDKEW =
UQYETKFQ-DKAPE =
EPX =
HEEQRMEKXG =
TXRQCPQG =
WQTEQDCQQB =
EHUUJXRKGAT =

Explanation

Try the questions first. Use the Hint button next to any part if you need a nudge.

Reading the explanation costs no XP. Worth 20 points in the real exam.