Information Retrieval

Programming, Math · course

Information Retrieval – 4

Michel Beigbeder

October, 13th 2016

– Typeset by FoilTEX –

IR Model

indexation

doc

dn

txt

d1

ps

d2

pdf

d3

corpus

δ1

δ2

δ3

δn

base d’index

U?

requˆete q

utilisateur

´evaluation

f (q, δi)

@dr(q,1)

@dr(q,2)

@dr(q,3)

mise en correspondance

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

1

TREC Model of IR (1/2)

Test collection : documents, information needs, relevance judgments

doc file

txt

d1

txt

d2

txt

d3

➊ corpus

filtre

➋

indexation

➍ besoins d’informations

➎ requˆetes q

✍

δ1

δ2

δ3

δn

➌ base d’index

➏

f (q, δi)

mise en correspondance

➑ jugements de pertinence

➐ runq

@dr(q,1)

@dr(q,2)

@dr(q,3)

´evaluation

➒ pr´ecision

rappel

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

2

TREC Model of IR (2/2)

➊ documents

➋ collection files

➌ documents’ index

➍ information needs

➎ queries

➏ queries’ index

➐ answer list as returned by the information retrie-

val system (run)

➑ relevance judgments

➒ evaluation (recall-precision)

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

3

One document from collection adi

Extract of file ➊ adi/adi.all :

Extract of its ➌ smart’s index :

docid concept_type concept_id weight term

[...]

17 0

.I 17

17 0

.T

17 0

document recovery

17 0

.A

17 0

R. L. BIRCG

17 0

.W

17 0

the naming of journals and organizations : implications for

17 0

names are chosen for technical journals for societies often

17 0

incorporating factors which needlessly complicate filing and

17 0

recovery . changes of name also appear to often ignore the

effect on commonplace information retrieval procedures .

17 0

factors considered include ease of memory retention,

17 0

compatibility of wording and of layout of title pages with

17 0

filing systems used in offices, libraries, and bibliographies .

17 0

.I 18

17 0

.T

17 0

state-of-the-art : remote interrogation of stored documentary

17 0

material

17 0

.A

17 0

H. OHLMAN

17 0

[...]

[...]

2.00000

887

1.00000

1061

1.00000

1763

1.00000

1877

1.00000

2802

1.00000

4112

1.00000

5991

1.00000

8143

1.00000

9309

9654

1.00000

10370 1.00000

12640 1.00000

13631 1.00000

17726 1.00000

18494 1.00000

18943 1.00000

19520 2.00000

19903 2.00000

19911 1.00000

20443 1.00000

the

implications

offices

effect

chosen

name

layout

information

changes

pages

organizations

also

incorporating

title

in

retention

filing

recovery

ease

technical

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

4

First information need for adi

Beginning of file ➍ adi/query.text :

.I 1

.W

What problems and concerns are there in making up descriptive titles?

What difficulties are involved in automatically retrieving articles from approximate titles?

What is the usual relevance of the content of articles to their titles?

[...]

Relevance judgments

Extract from file

➑ adi/qrels.text :

Same information with current

trec eval syntax :

qid docid

17

1

46

1

62

1

12

2

2

71

[...]

0

0

0

0

0

0.000000

0.000000

0.000000

0.000000

0.000000

qid

1

1

1

2

2

[...]

unused

0

0

0

0

0

docid relevance level

17

46

62

12

71

1

1

1

1

1

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

5

Three experiments on collection adi

Three runs :

➐ smart.nnn.nnn

➐ smart.lic.ann

➐ zettair

qid

1

1

1

1

1

1

1

1

1

...

0

0

0

0

0

0

0

0

0

docid

16

36

1

28

58

9

24

46

15

0

0

0

0

0

0

0

0

0

score

56.0000

41.0000

39.0000

39.0000

39.0000

38.0000

37.0000

37.0000

34.0000

run name

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

nnn.nnn

qid

Publicité

1

1

1

1

1

1

1

1

1

...

0

0

0

0

0

0

0

0

0

docid

69

27

47

30

19

25

37

22

46

0

0

0

0

0

0

0

0

0

score

0.4780

0.4526

0.3195

0.2820

0.2744

0.2508

0.2406

0.2305

0.2239

run name

lic.ann

lic.ann

lic.ann

lic.ann

lic.ann

lic.ann

lic.ann

lic.ann

lic.ann

qid

1

1

2

2

2

2

2

2

2

...

0

0

0

0

0

0

0

0

0

docid

69

46

71

69

68

35

75

64

23

0

0

0

0

0

0

0

0

0

score

4.658154

3.451871

5.410110

4.658154

4.365011

3.736974

3.701425

3.615445

3.478283

run name

zettair

zettair

zettair

zettair

zettair

zettair

zettair

zettair

zettair

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

6

Evaluation of an IRS (Information Retrieval System)

Assuming that

— relevance is binary

— the IRS is an input-output box (no interaction)

evaluation of the IRS ability :

— to return the relevant documents ;

— to NOT return the NON relevant documents.

Precision and recall

Recall =

|Rel ∩ Retr |

|Rel |

Precision =

|Rel ∩ Retr |

|Retr |

An IRS retrieves documents ➐.

Relevance judgments indicate what are the relevant documents ➑.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

7

F-measure

Weighted harmonic mean of P and R

F =

1

P + (1 − α) 1

α 1

R

F =

(β2 + 1)P R

β2P + R

avec β2 =

1 − α

α

In particular :

F1 = Fβ=1 =

2P R

P + R

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

8

Precision-Recall curves

In the ranked list, Recall and Precision are evaluated at each

rank by considering the set of retrieved documents up to this

rank.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

9

Query 27

33 relevant documents

66 retrieved documents sorted by decreasing score

P= 0/1 = 0.0%

P= 1/2 =50.0%

P= 2/3 =66.7%

P= 2/4 =50.0%

P= 2/5 =40.0%

P= 3/6 =50.0%

P= 3/7 =42.9%

P= 4/8 =50.0%

P= 5/9 =55.6%

P= 6/10=60.0%

P= 6/11=54.5%

P= 6/12=50.0%

P= 7/13=53.8%

P= 7/14=50.0%

P= 8/15=53.3%

P= 9/16=56.2%

P= 9/17=52.9%

P=10/18=55.6%

P=11/19=57.9%

P=12/20=60.0%

P=13/21=61.9%

R= 0/33= 0.0%

R= 1/33= 3.0%

R= 2/33= 6.1%

R= 2/33= 6.1%

R= 2/33= 6.1%

R= 3/33= 9.1%

R= 3/33= 9.1%

R= 4/33=12.1%

R= 5/33=15.2%

R= 6/33=18.2%

R= 6/33=18.2%

R= 6/33=18.2%

R= 7/33=21.2%

R= 7/33=21.2%

R= 8/33=24.2%

R= 9/33=27.3%

R= 9/33=27.3%

R=10/33=30.3%

R=11/33=33.3%

R=12/33=36.4%

R=13/33=39.4%

65

48

30

58

67

22

28

61

11

2

52

43

20

50

8

41

17

70

66

57

6

−

+

+

−

−

+

−

+

+

+

−

−

+

−

+

+

−

+

+

+

+

...

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

27

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

22

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

10

Query 27

33 relevant documents

66 retrieved documents sorted by decreasing score

65

48

30

58

Publicité

67

22

28

61

11

2

52

43

20

50

8

41

17

70

66

57

6

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

R= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

P= / = %

−

+

+

−

−

+

−

+

+

+

−

−

+

−

+

+

−

+

+

+

+

...

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

27

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

22

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

11

Interpolation - Extrapolation

COLL=adi.qrels RUN=zettair.adi.Q1.run

To be able to average the

values for different queries,

we want do define a value

for each recall point :

— interpolation

bet-

ween two points

to-

— extrapolation

wards R=0%

11 standard recall points :

from 0% to 100% by 10%

steps.

)

P

(

i

i

n

o

s

c

e

r

p

27

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

22

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

12

Average over the topics

COLL=adi.qrels RUN=zettair.adi.Q1.run

0.7

0.6

0.5

0.4

0.3

0.2

0.1

0

0

tr.nnn.nnn

tr.lic.ann

zettair.adi.Q1

0.2

0.4

0.6

0.8

1

)

P

(

i

i

n

o

s

c

e

r

p

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

COLL=adi.qrels RUN=zettair.adi.Q1.run

)

P

(

i

i

n

o

s

c

e

r

p

100

80

60

40

20

0

0 20 40 60 80 100

recall (R)

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

13

trec eval

— the standard tool for IR evaluation

— http://trec.nist.gov/trec_eval/trec_eval.8.1.tar.gz

— provides many measures :

— precision at 5, 10, . . ., 1000 documents

— R-Precision : precision at R (R is the number of relevant

documents for the topic)

— bpref

— Reciprocal rank

— . . .

— http://trec.nist.gov/pubs/trec16/appendices/measures.pdf

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

14

MAP : Mean Average Precision

AP(q) =

1

|Relq|

|Relq|

Xk=1

P (Rqk)

MAP(Q) =

1

|Q| Xq∈Q

AP(q)

P (Rqk) is precision at the rank of the k-th relevant retrieved docu-

ment.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

15

Some historic test collections (1/2)

Available at ftp://ftp.cs.cornell.edu/pub/smart/.

— adi.all, 82 abstracts of the articles of the American Do-

cumentation Institute 1963’s meeting, information science

domain.

— cacm.all.Z, 3 204 documents only containing title and cita-

tion links, computer science domain.

— cisi.all.Z, information science domain.

— med.all.Z, medical domain.

— npl.dat.Z, electronics, computer science and physics.

— time/doc.text.Z, news.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

16

Some historic test collections (2/2)

cranfield

adi

cacm

cisi

cran

med

npl

time

1398 docs

Publicité

36 K

82 docs title

2.1 M

3204 docs title

2.3 M

1460 docs title

1.6 M

1400 docs title

1033 docs

1.0 M

3.1 M 11429 docs

425 docs

1.5 M

authors

authors

authors

authors

abstract

abstract

abstract

abstract

abstract

long title

news

citations

citations

225 requ.

35 requ.

64 requ.

112 requ.

225 requ.

30 requ.

93 requ.

83 requ.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

17

TREC : Text Retrieval conference

TREC is the most well known annual IR evaluation campaign.

It has been existing since 1992. 23rd edition in 2014. The cycle of

the campaign :

— NIST proposes tasks and provides the sets of documents and

≪topics ≫ to the participants ;

— each participant runs its system with these data and builds

an answer list of results for each topic (1 000 top documents)

(RUN ) ;

— evaluation of the runs by NIST

— november workshop at NIST.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

18

Some TREC tracks

— Adhoc : 1992–1999 Robust : 2003–2005

— Interactive : 1994–2005

— Routing : 1992–1997 Filtering : 1995–2002 Spam :2005–

— Spanish : 1994–1996 Chinese : 1996–1997 Multilingual : 1997–

2002

— OCR, speech, video (now VidTREC)

— Query Answering : 1995–2007

— By domains : Legal, Web, Chemical, Medical, e-mail, etc.

— etc. etc.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

19

Questions about test collections

— Criteria for choosing the documents ?

— Task representativity

— Diversity of subjets, of the vocabulary

— Data : Full text vs. abstract, hyperlinks, social data,. . .

— . . .

— How many topics ?

— How to identify the relevant documents for each topic ?

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

20

TREC collections characteristics

Each track uses its own topics (usually fifty ones).

— Documents from newspapers

Associate Press Newswire (1988–1989)

WSJ Wall Street Journal (1986–1992)

AP

ZIFF Ziff-Davis Publishing

Federal Register (1988–1989)

FR

DOE abstracts from Dept. Energy reports

550 M 173252 docs

514 M 164597 docs

493 M 132100 docs

469 M 45820 docs

190 M 226087 docs

— Documents from the Web

1997

VLC

extrait de VLC

WT2G

WT10G extrait de VLC

.GOV

W3C

extrait des sites .gov (2003) HTML, PDF, etc.

le site du W3C (2004)

HTML, PDF, etc.

— Medline

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

21

HTML

HTML

HTML

100 G

2 G

10 G

500 G

? G

A sample document

<DOC>

<DOCNO> WSJ870324-0001 </DOCNO>

<HL> John Blair Is Near Accord To Sell Unit, Sources Say </HL>

<DD> 03/24/87</DD>

<SO> WALL STREET JOURNAL (J) </SO>

<IN> REL TENDER OFFERS, MERGERS, ACQUISITIONS (TNM)

MARKETING, ADVERTISING (MKT) TELECOMMUNICATIONS,

BROADCASTING, TELEPHONE, TELEGRAPH (TEL) </IN>

<DATELINE> NEW YORK </DATELINE>

<TEXT>

John Blair & Co. is close to an agreement to sell its TV station

advertising representation operation and program production unit to an

investor group led by James H. Rosenfield, a former CBS Inc.

executive, industry sources said. Industry sources put the value of the

proposed acquisition at more than $100 million. ...

</TEXT>

</DOC>

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

22

Some statistics on the collections

CACM

CISI

WSJ-1

AP-1

ZIFF-1

FR-1

DOE

WSJ-2

AP-2

ZIFF-2

FR-2

WT10G

GOV2

ClueWeb09

2 M

2 M

267 M

254 M

242 M

260 M

184 M

242 M

237 M

175 M

209 M

10 G

500 G

25 T

3 204 docs

1 460 docs

98 732 docs

84 678 docs

75 180 docs

25 960 docs

226 087 docs

74 520 docs

79 919 docs

56 920 docs

19 860 docs

1 692 096 docs

25 205 179 docs

1 040 809 705 docs

tokens/doc

median

245

446

200

391

111

301

438

182

396

tokens/doc

mean

40.1

104.9

434.0

473.9

473.0

1315.9

120.4

508.4

468.7

451.9

1378.1

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

23

What is a Topic ?

— Description of an information need

— Gives some indications about relevance

— Build by an assessor

— The assessor who built the topic will

documents for this topic

identify the relevant

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

24

A sample topic

<top>

<num> Number: 451

<title> What is a Bengals cat?

<desc> Description:

Provide information on the Bengal cat breed.

<narr> Narrative:

Item should include any information on the Bengal cat breed,

including description, origin, characteristics, breeding program,

names of breeders and catteries carrying bengals.

References which discuss bengal clubs only are not relevant.

Discussions of bengal tigers are not relevant.

</top>

NB : The query has to be built on the basis of the topic.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

25

Questions about the topics

— How many topics ?

— Impact of the number of topics (C. Buckley et E. Voorhees

[SIGIR 2000])

— the number of topics must be greater than 25

— 50 seems a good value

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

26

Relevance judgments

How to identify the relevant documents for each topic ? Spending 30 s to

judge a document, it would need 6 500 hours to judge 800 000 documents (more

than 900 working days) So TREC uses the pooling method.

Questions about relevance judgments

— Consistency

— Relevance is subjective and depends on the assessor.

— What happens if changing the assessors ?

— Study of E. Voorhees [IPM 2000]

— Completion

— Some documents ARE relevant but NOT judged

— These documents are considered NON relevant

— Are the systems which did not contribute to the pooling penalized ?

— Study of Zobel [SIGIR 1998]

— bpref

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

27

bpref

bpref q =

1

R

Xr (cid:18)

1 −

|n ranked higher than r|

min(R, N )

(cid:19)

where

— R is the number of judged relevant documents,

— N is the number of judged irrelevant documents,

— r is a relevant retrieved document,

— and n is a member of the first R irrelevant retrieved docu-

ments.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

28

Other initiatives

— NTCIR since 1997 on documents in asian languages (NII-

NACSIS Test Collection for IR Systems)

— CLEF, created in 2000, and dedicated to multilingual

IR

evaluation (fr, en, es, it, de, g, sw, fi, etc.)

— CLEF is now a more general evaluation forum (Patents, Me-

dical, Social, Cultural Heritage, Plagiarism,. . .)

— INEX began in 2002 for structured IR evaluation (XML docu-

ments). Now part of CLEF.

— FIRE since 2008, South Asian language Information Access.

EMSE M. Beigbeder

2017–18

Master DSC 2A

IR

29