Skip to main navigation Skip to search Skip to main content

Assessement of NER solutions against the first and second CALBC Silver Standard Corpus

  • D Rebholz-Schuhmann
  • , A Jimeno Yepes
  • , C Li
  • , S Kafkas
  • , I Lewin
  • , Ning Kang
  • , P Corbett
  • , D Milward
  • , E Buyko
  • , E Beisswanger
  • , K Hornbostel
  • , A Kouznetsov
  • , R Witte
  • , JB Laurila
  • , CJO Baker
  • , C-J Kuo
  • , S Clematide
  • , F Rinaldi
  • , R Farkas
  • , G Mora
  • K Hara, LI Furlong, M Rautschka, M Lara Neves, A Pascual-Montano, Q Wei, N Collier, FM Chowdhury, A Lavelli, R Berlanga, R Morante, V van Asch, W Daelemans, J Luis Marina, Erik van Mulligen, Jan Kors, U Hahn

Research output: Contribution to journalArticleAcademic

43 Citations (Scopus)
3 Downloads (Pure)

Abstract

Competitions in text mining have been used to measure the performance of automatic text processing solutions against a manually annotated gold standard corpus (GSC). The preparation of the GSC is time-consuming and costly and the final corpus consists at the most of a few thousand documents annotated with a limited set of semantic groups. To overcome these shortcomings, the CALBC project partners (PPs) have produced a large-scale annotated biomedical corpus with four different semantic groups through the harmonisation of annotations from automatic text mining solutions, the first version of the Silver Standard Corpus (SSC-I). The four semantic groups are chemical entities and drugs (CHED), genes and proteins (PRGE), diseases and disorders (DISO) and species (SPE). This corpus has been used for the First CALBC Challenge asking the participants to annotate the corpus with their text processing solutions.
The SSC-I delivers a large set of annotations (1,121,705) for a large number of documents (100,000 Medline abstracts). The annotations cover four different semantic groups and are sufficiently homogeneous to be reproduced with a trained classifier leading to an average F-measure of 85%. Benchmarking the annotation solutions against the SSC-II leads to better performance for the CPs’ annotation solutions in comparison to the SSC-I.

Original languageEnglish
Pages (from-to)11
JournalJournal of Biomedical Semantics
Volume2
Issue numbersupplement
DOIs
Publication statusPublished - 6 Oct 2011

Research programs

  • EMC NIHES-03-77-01

Fingerprint

Dive into the research topics of 'Assessement of NER solutions against the first and second CALBC Silver Standard Corpus'. Together they form a unique fingerprint.

Cite this