Please share with graduate and undergraduate students looking for summer internships. The Johns Hopkins University Human Language Technology Center of Excellence (COE) is seeking candidates for our summer internship program as part of SCALE, our yearly summer workshop (Summer Camp for Advanced Language Exploration.) Interns will work on research in speech, text and graph processing as part of a larger workshop team. Internship positions are fully funded, including travel, living expenses and stipend. This summer's workshop topic is "Combining Speech, Text and Graphs," which will draw from a number of research areas: *** Natural Language Processing *** - machine translation and NLP on informal texts - opinion extraction from informal texts - topic models - information extraction - domain adaptation and transfer learning - multi-lingual learning *** Speech *** - topic detection - keyword spotting - resource constrained speech processing - robust speech processing *** Graphs *** - finding communities in social networks - anomaly detection - learning on large scale graphs - graph-based topic modeling Candidates should be currently enrolled in an undergraduate or graduate degree program. Applications submitted by Friday Jan 14, 2011 will receive full consideration. For more information: http://web.jhu.edu/HLTCOE/scaleinterns2011.html Applicants will be required to obtain a US security clearance, which requires US citizenship. If you do not already have a clearance, we will work with you to obtain one.
22 December 2010
Summer internships at JHU/COE
... in case anyone is reading this, I just got the following email. I know this program and it's good.
21 December 2010
Grades (Almost) Posted, Semester Over
Hi all --
I figure you're likely to read to the end to find out about grades, so before I get to that, let me just take this chance to say that I really enjoyed this class this semester. You were all great. Everyone did awesome on both the final exam and final projects, and I'm really thrilled. If you couldn't tell already, I love language stuff and I encourage you all to continue on and learn everything there is to know. Philip is teaching the follow-on course in the Spring, which should be awesome. I'm also running an unofficial seminar on "Large Data" stuff in the spring; you can get more info here (sign up for the mailing list if you're interested). Anyway, I had a great time teaching; I hope you had fun in class.
Regarding grades, I need to submit them by midnight tonight. And since I don't plan on staying up until midnight, this really means tonight by about 11p.
I've posted "unofficial" grades on grades.cs.umd.edu, so you can see what your grade is. Note that the "total" column on that spreadsheet is completely misleading, since it doesn't include all the weirdo grading rules (dropping of worst projects/homeworks, inclusion of extra credit, etc.). I have all the numbers in a separate spreadsheet, so if something looks odd to you and you'd like the full computation, please let me know. It's of course possible to change grades later, but it's a pain, so I'd rather hear about any issues now.
That's it. Have a great break and I hope to see some of you again in the Spring!
-h
ps., per second comment below, I added an extra column, MyOverall. The grading is as follows:
98 = A+
95 = A
92 = A-
88 = B+
85 = B
82 = B-
78 = C+
75 = C
72 = C-
Note that your score will be exactly one of these numbers: This is just my way of encoding your grade. This isn't actually what your score was :).
I figure you're likely to read to the end to find out about grades, so before I get to that, let me just take this chance to say that I really enjoyed this class this semester. You were all great. Everyone did awesome on both the final exam and final projects, and I'm really thrilled. If you couldn't tell already, I love language stuff and I encourage you all to continue on and learn everything there is to know. Philip is teaching the follow-on course in the Spring, which should be awesome. I'm also running an unofficial seminar on "Large Data" stuff in the spring; you can get more info here (sign up for the mailing list if you're interested). Anyway, I had a great time teaching; I hope you had fun in class.
Regarding grades, I need to submit them by midnight tonight. And since I don't plan on staying up until midnight, this really means tonight by about 11p.
I've posted "unofficial" grades on grades.cs.umd.edu, so you can see what your grade is. Note that the "total" column on that spreadsheet is completely misleading, since it doesn't include all the weirdo grading rules (dropping of worst projects/homeworks, inclusion of extra credit, etc.). I have all the numbers in a separate spreadsheet, so if something looks odd to you and you'd like the full computation, please let me know. It's of course possible to change grades later, but it's a pain, so I'd rather hear about any issues now.
That's it. Have a great break and I hope to see some of you again in the Spring!
-h
ps., per second comment below, I added an extra column, MyOverall. The grading is as follows:
98 = A+
95 = A
92 = A-
88 = B+
85 = B
82 = B-
78 = C+
75 = C
72 = C-
Note that your score will be exactly one of these numbers: This is just my way of encoding your grade. This isn't actually what your score was :).
14 December 2010
12 December 2010
Interested in Language Science?
Just got the following email from Colin Philips in Linguistics / Cognitive Neuroscience. This is regarding language science. Please see below... feel free to email me if you have questions:
I'm hoping that you can help us to reach students in the CS/iSchool universe who might be interested in taking advantage of our unique interdisciplinary language science opportunities. We're particularly interested in reaching 1st and 2nd year students. We'll be holding an informational meeting for students tomorrow at 1pm in 1108B Marie Mount, but I'd be happy to meet at another time with anybody who is interested but not available at that time. We'll tell students about the opportunities and benefits, and also talk about the resources that are available to help them, including new plans to help them to develop interdisciplinary training plans that are both innovative and feasible. Csilla Kajtar already circulated a message about this, but we know that people often just ignore messages sent to mailing lists.
As you know, the closer integration of NLP and cognitive work in language is right at the top of our list of opportunities-that-we'd-be-idiots-not-to-pursue, and student training is one of the best ways to achieve this.
09 December 2010
Final Exam, Due Dec 17, 3:30pm
Here's a copy of the final exam as well as the source LaTeX. Please feel free to either print it and do it by hand, or to do it in LaTeX and print the solution. You may turn it in any time between now and 3:30pm on Dec 17. (Because our official exam time is 1:30-3:30 on Dec 17.) Please hand it in in one of three ways: (1) give it to me in person in my office or otherwise; (2) slide it completely under my office door (AVW 3227); (3) give it to Amit in person.
If you have any clarification questions, please post them here.
If you have any clarification questions, please post them here.
06 December 2010
P4 deadline pushed back to Dec 14
The 9th is apparently the deadline for the ML project.
05 December 2010
P4: Small error in example numbers....
There are some sentences in the training data that contain a "TAB" character. The reasonable thing to do would be just to consider this as whitespace. For some reason I didn't do this. In my example of DF computation, I did this. Which somewhat changes all the remaining numbers.
Instead of rerunning everything I'll just tell you what the updated top frequency words are if you do it "properly." In general, for this assignment, don't worry if your numbers are slightly different than mine -- it may have to do with how you handle the non-ascii characters that appear once in a while in the data.
Instead of rerunning everything I'll just tell you what the updated top frequency words are if you do it "properly." In general, for this assignment, don't worry if your numbers are slightly different than mine -- it may have to do with how you handle the non-ascii characters that appear once in a while in the data.
2999 . 2999 , 2998 of 2997 the 2997 and 2994 in 2989 to 2988 a 2885 as 2875 by 2862 for 2860 ) 2859 ( 2836 with 2832 that 2801 '' 2788 `` 2759 on 2717 from
Last seminar of the semester: Michael Paul Dec 8, 11am
December 8: Michael Paul: Summarizing Contrastive Viewpoints in Opinionated Text
AVW 2120 Performing multi-document summarization of opinionated text has unique challenges because it is important to recognize that the same information may be presented in different ways from different viewpoints. In this talk, we will present a special kind of contrastive summarization approach intended to highlight this phenomenon and to help users digest conflicting opinions. To do this, we introduce a new graph-based algorithm, Comparative LexRank, to score sentences in a summary based on a combination of both representativeness of the collection and comparability between opposing viewpoints. We then address the issue of how to automatically discover and extract viewpoints from unlabeled text, and we experiment with a novel two-dimensional topic model for the task of unsupervised clustering of documents by viewpoint. Finally, we discuss how these two stages can be combined to both automatically extract and summarize viewpoints in an interesting way. Results are presented on two political opinion data sets. This project was joint work with ChengXiang Zhai and Roxana Girju. Bio: Michael Paul is a first-year Ph.D. student of Computer Science at the Johns Hopkins University and a member of the Center for Language and Speech Processing. He earned a B.S. from the University of Illinois at Urbana-Champaign in 2009. He is currently a Graduate Research Fellow of the National Science Foundation and a Dean's Fellow of the Whiting School of Engineering.
02 December 2010
Lecture 25: Mapping Text to Actions
There has been a bunch of work recently on trying to automatically find relationships between language and the "real world", where "real world" actually often means some sort of simulated environment. Here are a few papers along these lines:
In the first paper, which is the one we'll talk about most, the key idea is that of hierarchical plans, represented as a pcfg. For instance we might have a rule "OfferCup -> PickUpCup MoveCup ReleaseCup", where each of the subactions might either be atomic (correspond to actual muscle movements) or might itself be broken down further. (Qustion: how context free is this problem?)
The key ambiguity is due to the fact that actions do not select for exactly one interpretation, as in the Blicket example.
In this paper, they hand constructed a PCFG for actions and the key learning question was whether you could figure out the level of ambiguity automatically. The basic idea is to look at relative frequencies of occurance between lexical items and nodes in the PCFG tree for the actions.
- Fleischman, M. B. and Roy, D. Intentional Context in Situated Language Learning. Ninth Conference on Computational Natural Language Learning , Ann Arbor, MI. June 2005.
- Learning to Connect Language and Perception [Abstract] [PDF]
Raymond J. Mooney
In Proceedings of the 23rd AAAI Conference on Artificial Intelligence (AAAI), Senior Member Paper, Chicago, IL, pp. 1598-1601, July 2008. - S.R.K. Branavan, Harr Chen, Luke Zettlemoyer and Regina Barzilay
"Reinforcement Learning for Mapping Instructions to Actions",
Proceedings of ACL, 2009. Best Paper Award - Learning semantic correspondences with less supervision.
Percy Liang, Michael I. Jordan, Dan Klein.
Association for Computational Linguistics and International Joint Conference on Natural Language Processing (ACL-IJCNLP), 2009. - Adam Vogel and Dan Jurafsky. 2010. Learning to Follow Navigational Directions. In Proceedings of ACL-2010, Uppsala, Sweden. [PDF]
In the first paper, which is the one we'll talk about most, the key idea is that of hierarchical plans, represented as a pcfg. For instance we might have a rule "OfferCup -> PickUpCup MoveCup ReleaseCup", where each of the subactions might either be atomic (correspond to actual muscle movements) or might itself be broken down further. (Qustion: how context free is this problem?)
The key ambiguity is due to the fact that actions do not select for exactly one interpretation, as in the Blicket example.
In this paper, they hand constructed a PCFG for actions and the key learning question was whether you could figure out the level of ambiguity automatically. The basic idea is to look at relative frequencies of occurance between lexical items and nodes in the PCFG tree for the actions.
01 December 2010
P4, grading rules
So P4 has been posted for a while. It is "optional" in the sense that your project grades will be based on your best three out of four grades. In particular, here's what will happen.
Suppose that your grades on P1, ..., P4 are a,b,c,d. (If you don't do P4, then d=0.)
Let x = [ (a + b + c + d) - min { a, b, c, d } ] / 3
Then x is your average grade on your best three projects.
We will use x as your overall project grade (i.e., since each project is weighed equally, it will be like you got a score of x on all FOUR of them).
Suppose that your grades on P1, ..., P4 are a,b,c,d. (If you don't do P4, then d=0.)
Let x = [ (a + b + c + d) - min { a, b, c, d } ] / 3
Then x is your average grade on your best three projects.
We will use x as your overall project grade (i.e., since each project is weighed equally, it will be like you got a score of x on all FOUR of them).
30 November 2010
Lecture 24: Information Extraction
Information extraction is, roughly, the task of going from unstructured text (aka text) to structured data. Think of it as mapping language to a database.
One of the more famous IE tasks is identifying terrorist events (specifically South American terrorist events) in documents. For each event, we have to identify the victim(s), date, type of event (bombing, etc.), culprits, and so on. These fields define our extraction template and our goal is to fill it up based on a document. Of course some documents mention no terrorist events and some mention multiple. And not all fields will be mentioned. This data was available for the MUC (message understanding conference) competitions two decades ago and current approaches still only get about 50% accuracy!
One way of going about this problem is as a sequence labeling task, akin to NER. Since you can imagine how this works, we'll talk about the other style of approach: pattern-based methods.
The idea of pattern-based approaches is that our system consists of a collection of extraction patterns, of the form "<Subj> was assassinated" => Victim. These lexico-syntactic patterns tell us how and when to extract certain slots (aka fields) from text. The expressiveness of patterns depends entirely on how much preprocessing you want to do, but usually some sort of syntactic processing is assumed.
The key question is: where do these patterns come from? The trend in IE has been to move toward mostly unsupervised approaches that don't need large amounts of training data. Successful approaches are akin to the Yarowsky algorithm for WSD.
Suppose we had labeled data, where for each document we have one (for simplicity) partially filled template. We can go into the document and find occurrences of the strings in that template as the things we want to extract. For each of these, we can find a bunch of extraction patterns that would potentially extract that string and collect them over the whole data set. We now need to find the "best" ones. A common metric is the "R log F" metric, which is simply the probability that a given pattern extracts the right slot, times log of the frequency of that pattern. The "log F" term is in there because we want to make sure that we get good coverage.
Of course, you needn't start with labeled data. You can start with small lists of slot fillers (eg., Al Queda as a perpetrator and so on) and bootstrap away. As always, the quality of your seeds directly affects how well your algorithm works.
One can get even more unsupervised by doing the following. Take a collection of documents that talk about terrorist events, and a collection of documents that don't. Look for patterns in the terrorist events collection that are significantly more common there, than in the other collection. Rank these by something like "R log F". The top patterns there are often very good extraction patterns, but we don't know what they are supposed to extract. Have a human look down the list of the top 100 and viola, you're done, and it only takes a few minutes.
Most approaches to IE fall into one of these two camps: sequence labeling or pattern based approaches. It seems that sequence labeling approaches work well when most of the words in the text are extracted for something (i.e., turning free text citations into bibtex entries), but pattern based approaches work well for needle in a haystack problems.
There have been a few recent trends in IE:
One of the more famous IE tasks is identifying terrorist events (specifically South American terrorist events) in documents. For each event, we have to identify the victim(s), date, type of event (bombing, etc.), culprits, and so on. These fields define our extraction template and our goal is to fill it up based on a document. Of course some documents mention no terrorist events and some mention multiple. And not all fields will be mentioned. This data was available for the MUC (message understanding conference) competitions two decades ago and current approaches still only get about 50% accuracy!
One way of going about this problem is as a sequence labeling task, akin to NER. Since you can imagine how this works, we'll talk about the other style of approach: pattern-based methods.
The idea of pattern-based approaches is that our system consists of a collection of extraction patterns, of the form "<Subj> was assassinated" => Victim. These lexico-syntactic patterns tell us how and when to extract certain slots (aka fields) from text. The expressiveness of patterns depends entirely on how much preprocessing you want to do, but usually some sort of syntactic processing is assumed.
The key question is: where do these patterns come from? The trend in IE has been to move toward mostly unsupervised approaches that don't need large amounts of training data. Successful approaches are akin to the Yarowsky algorithm for WSD.
Suppose we had labeled data, where for each document we have one (for simplicity) partially filled template. We can go into the document and find occurrences of the strings in that template as the things we want to extract. For each of these, we can find a bunch of extraction patterns that would potentially extract that string and collect them over the whole data set. We now need to find the "best" ones. A common metric is the "R log F" metric, which is simply the probability that a given pattern extracts the right slot, times log of the frequency of that pattern. The "log F" term is in there because we want to make sure that we get good coverage.
Of course, you needn't start with labeled data. You can start with small lists of slot fillers (eg., Al Queda as a perpetrator and so on) and bootstrap away. As always, the quality of your seeds directly affects how well your algorithm works.
One can get even more unsupervised by doing the following. Take a collection of documents that talk about terrorist events, and a collection of documents that don't. Look for patterns in the terrorist events collection that are significantly more common there, than in the other collection. Rank these by something like "R log F". The top patterns there are often very good extraction patterns, but we don't know what they are supposed to extract. Have a human look down the list of the top 100 and viola, you're done, and it only takes a few minutes.
Most approaches to IE fall into one of these two camps: sequence labeling or pattern based approaches. It seems that sequence labeling approaches work well when most of the words in the text are extracted for something (i.e., turning free text citations into bibtex entries), but pattern based approaches work well for needle in a haystack problems.
There have been a few recent trends in IE:
- Using Wikipedia infoboxes as training data
- Trying to extract knowledge without pre-defined templates (akin to our discussion of mining world knowledge)
23 November 2010
Lecture 23: Rhetorical Structure Theory
So far we've seen flat representations of discourse. RST is an example of a hierarchical discourse representation. Such as:
Here, we've broken some imagined text into 7 "units" and depicted the role of these units in the text. As is implied by the title, this theory is mostly applicable to rhetoric, which is essentially persuasive language.
An RST structure is essentially a dependency tree over elementary discourse units (EDUs), where relations on the edges of the tree tell us the relationship between two EDUs. In the above example, we're saying that EDUs 1-3 provide background to EDUs 4-7. And so on down the tree.
For most relations, there is a distinction between the nucleus and satellite of that relation: basically this is just like headedness in syntax. The nucleus contains the important stuff.
The relations in RST are given by communicative intent. This is a big separation between RST and other theories of discourse.
Here is an example of the evidence relation:
For instance, concession is often identified by the word "although." And "Evidence" is often identified by "For instance." And "Elaboration" is often identified by "And." And so on.
One clever idea a few years ago was to try to mine lexical relations that are indicative of discourse structure. For example, I can find all sentences that begin "for example" and look at that sentence, and the preceding sentence. I can assume that this is an example of Evidence, and then look at features of those two sentences to try to figure out why this is an evidence relation. Then, in the future, when I see sentences that don't have this lexical cue, I can apply whatever I've learned.
The hope is that if you mine contrast relations, you can find contrasting pairs like love/hate or iPhone/Android or whatever. As was shown in the assigned paper for today, that didn't work particularly well. My feeling is that lexical information is not "deep" enough to really get you to discourse except in very simple cases (see that post I linked to before).
Here, we've broken some imagined text into 7 "units" and depicted the role of these units in the text. As is implied by the title, this theory is mostly applicable to rhetoric, which is essentially persuasive language.
An RST structure is essentially a dependency tree over elementary discourse units (EDUs), where relations on the edges of the tree tell us the relationship between two EDUs. In the above example, we're saying that EDUs 1-3 provide background to EDUs 4-7. And so on down the tree.
For most relations, there is a distinction between the nucleus and satellite of that relation: basically this is just like headedness in syntax. The nucleus contains the important stuff.
The relations in RST are given by communicative intent. This is a big separation between RST and other theories of discourse.
Here is an example of the evidence relation:
- The program as published for the calendar year 1980 really works.
- In only a few minutes, I entered all the figures from my 1980 tax return and got a result which agreed with my hand calculations to the penny.
- constraints on N: R might not believe N to a degree satisfactory to W
- constraints on S: R believes S or will find it credible
- constrains on N+S: R's comprehending S increases R's belief of N
- the effect: R's belief of N is increased
- locus of effect: N
- Concern that this material is harmful to health or the environment may be misplaced.
- Although it is toxic to certain animals,
- evidence is lacking that it has any serious long-term effect on human beings.
- constraints on N: W has positive regard for the situation presented in N
- constraints on S: W is not claiming that the situation presented in S doesn't hold
- constraints on N+S: W acknowledges a potential or apparent incompatibility between the situations presented in N and S; W regards the situations presented in N and S as compatible; recognizing that the compatibility between the situations presented in S and S increases R's positive regard for the situation presented in N
- the effect: R's positive regard for the situation presented in N is increased
- locus of effect: N and S
- Circumstance
- Solutionhood
- Elaboration
- Background
- Enablement and Motivation
- Evidence and Justify
- Relations of Cause
- Antithesis and Concession
- Condition and Otherwise
- Interpretation and Evaluation
- Restatement and Summary
- Sequence
- Contrast
For instance, concession is often identified by the word "although." And "Evidence" is often identified by "For instance." And "Elaboration" is often identified by "And." And so on.
One clever idea a few years ago was to try to mine lexical relations that are indicative of discourse structure. For example, I can find all sentences that begin "for example" and look at that sentence, and the preceding sentence. I can assume that this is an example of Evidence, and then look at features of those two sentences to try to figure out why this is an evidence relation. Then, in the future, when I see sentences that don't have this lexical cue, I can apply whatever I've learned.
The hope is that if you mine contrast relations, you can find contrasting pairs like love/hate or iPhone/Android or whatever. As was shown in the assigned paper for today, that didn't work particularly well. My feeling is that lexical information is not "deep" enough to really get you to discourse except in very simple cases (see that post I linked to before).
22 November 2010
P3 grading updated, deadline Wed 24th, 5p
Looks like I underestimated the difficulty of the gender classification. I've adjusted the scoring to be easier on you. The new scoring is:
- 35 < e < 37 : 10%
- 34 < e < 35 : 25%
- 33 < e < 34 : 32%
- 32 < e < 33 : 34%
- 31.5 < e < 32 : 36%
- 31 < e < 31.5 : 37%
- 30.5 < e < 31.0 : 38%
- 30 < e < 30.5 : 39%
- e < 30: 40%
18 November 2010
Lecture 22: Document Coherence
Documents are not just collections of sentences. Sentences serve purposes, and well-written document puts its sentences together well.
There are (at least) two notions of "well":
This is a lot like syntax/semantics. A document can be cohesive but still make no sense. This is like syntax. On the other hand, a document can be meaningful, but still not feel like the sentences go together properly. This is like semantics.
TextTiling is an algorithm/model for discovering cohesive subparts (despite what the original paper says). This is basically a partitioning problem. We're given a text as a sequence of sentences, and we want to break it into pieces, each of which is internally cohesive. The algorithm looks roughly like:
We then define the splits as the points where the similarities drop below some threshold.
One major (linguistic) weakness of this approach is its inability to handle synonymy, etc. When we talk about lexical repetition for cohesion, we also typically include things like "train / car / engine" even though their not the same word.
Argumentative Zoning is more of a coherence-type model (though it's not precisely clear that it's only in one of the two categories). It is primarily a model for research papers, where we talk about the role that each sentence plays in a paper. Like TextTiling, the model is also flat: a document is just a sequence of sentences and we're going to assign each sentence to a class. The three major classes are { Background, Other, Own }.
The original approach to AZ is to annotate data, and then train a classifier. The features used break down into a handful of categories:
There has been a lot of interesting work in this area since then, including things like sentiment analysis of citations, better features, better learning, etc...
There are (at least) two notions of "well":
- Coherence: This is what makes a text meaningful. This is the idea that the semantic interpretation of the sentences in a text should fit together into a larger picture.
- Cohesion: This is what makes a text hang together. This is the idea that sentences help you understand the relationship between what came before and what is coming later.
This is a lot like syntax/semantics. A document can be cohesive but still make no sense. This is like syntax. On the other hand, a document can be meaningful, but still not feel like the sentences go together properly. This is like semantics.
TextTiling is an algorithm/model for discovering cohesive subparts (despite what the original paper says). This is basically a partitioning problem. We're given a text as a sequence of sentences, and we want to break it into pieces, each of which is internally cohesive. The algorithm looks roughly like:
- For each sentence, collect all the word (stems) that occur in that sentence.
- Define the similarity between any two token sequences (subsequences of the text) as the cosine similarity between their stem vectors.
- For each position i in the text, compute the similarity between the block i-21:i-1 and i:i+20. (20 is arbitrary.)
We then define the splits as the points where the similarities drop below some threshold.
One major (linguistic) weakness of this approach is its inability to handle synonymy, etc. When we talk about lexical repetition for cohesion, we also typically include things like "train / car / engine" even though their not the same word.
Argumentative Zoning is more of a coherence-type model (though it's not precisely clear that it's only in one of the two categories). It is primarily a model for research papers, where we talk about the role that each sentence plays in a paper. Like TextTiling, the model is also flat: a document is just a sequence of sentences and we're going to assign each sentence to a class. The three major classes are { Background, Other, Own }.
The original approach to AZ is to annotate data, and then train a classifier. The features used break down into a handful of categories:
- explicit structure (where is this sentence in the document)
- relative location (what %age through the document is this sentence)
- citation (to self versus to others)
- syntax (tense, aspect, voice, negation, etc.)
- semantic (type of verb, etc.)
- content (word overlap with title, etc.)
There has been a lot of interesting work in this area since then, including things like sentiment analysis of citations, better features, better learning, etc...
16 November 2010
Lecture 21: Local Discourse Context
Coreference analysis means figuring out the meaning of pronouns in sentences like: "John saw him in the mirror" versus "John saw himself in the mirror." That's anaphor resolution. More generally, "John saw Bill and said 'My friend, you look lovely today.'" we need to figure out that My=John and friend=Bill.
Generally we care about three "mention" types:
There are lots of signals of coreference:
Generally we care about three "mention" types:
- Pronouns (he, she, it, I, etc...)
- Proper nouns (John, Bill, Congress, the White House, etc...)
- Common nouns (the soccer player, the man, etc.)
There are lots of signals of coreference:
- Apposition: "John, mayor of Tinyville, is happy." Here, it's pretty clear that John=mayor
- Syntactic structure with pronouns/-self. "John saw him" versus "... himself"
- Gender / number / type agreement. "He" is probably not "Mary" and "him" is probably not "players". Similarly, "he" is probably not "the White House". (Mr. and Mrs. are helpful here.)
- Locality. "He" probably refers back to a nearby referent, not something far away.
- Discourse focus. New entities are likely to transition from being objects to being subjects, not the other way around.
- Sub/super-string matching. "Bill Clinton" and "Clinton" are likely coreferent.
- New entities are often introduced with indeterminant NPs ("a player") and then later with determinant NPs
- World knowledge: sometimes we just know that Obama is the President
- Named to named: very easy, 95% accuracy just using substring matching
- Named to pronoun: pretty easy, 85-90% using Hobbs' heuristics, locality and gender
- Named to common: very hard (often need world knowledge)
- Common to common: very hard (often need lexical semantics)
- Common to pronoun: don't bother -- do common to named and named to pronoun
- You then need to go back and fix things up to make sure that transitivity is enforced; or use left-to-right search
- Number of -ve examples >> number of +ve examples; often subsample the negatives
- These decisions aren't independent
- Metonymy: when we say "the White House" meaning "Obama"
- Common nouns that aren't mentions of an entity: "This life is wonderful"
- Quantified pronouns: "Every man loves his mother" (might or might not be coreferent)
- World knowledge
15 November 2010
P2 is mostly graded...
If you'd like your partial score, please send me an email with the ID of the person who handed it in. I would love to just spam you all by sending your report to the hander-in at their id, but I don't know how to do that automatically. I tried sending an email to myid@umd.edu but it bounced... if anyone knows how to do that instead, let me know.
Overall people seemed to do quite well.
Overall people seemed to do quite well.
14 November 2010
TURKERS Final Project
Looks like there's no time tomorrow (Monday) that works for everyone. Please vote on your time preference, trying to be as flexible as possible, and figuring that we'll spend 30 minutes. Also, please list your email address in your signup so I can contact you! :)
11 November 2010
HW10 posted
and HW09 deadline pushed back to Sunday night
Subscribe to:
Posts (Atom)

