Showing posts with label bibliometrics. Show all posts
Showing posts with label bibliometrics. Show all posts

Friday, December 30, 2011

Correlation between reference managers and the WoS

Even though web citations have been a part of our lives for several years now, the correlation between "traditional" citations and web resources like Mendeley, CiteULike, blog networks, etc. hasn't been thoroughly studied yet, and any new research in the field is very interesting (to me, anyway).

The new paper was published at Scientometrics by Li, Thelwall (still one of my dissertation advisors) and Giustini. They focused on the correlation between user count - the number of users who save a particular paper - and WoS and Google Scholar citations.

The researchers extracted from WoS all the Nature and Science research articles that were published in 2007 and their references. They ended up with 793 Nature and 820 Science articles, or 1,613 articles overall (not including references, of course). Then, they searched CiteULike for those articles' titles and number of citations, as well as for their user count in Mendeley. They also collected the same data from Google Scholar. It's important to note that Mendeley had 32.9 million articles indexed while CiteULike had only 3.5 at the time of the study.

Google Scholar's mean and median number of citations were higher than in WoS (not surprising; If you want better citation numbers, always use GS). They found that despite Mendeley being "younger" than CiteULike (launched in 2004 and 2008 respectively), CiteULike had only about two-thirds of the sample articles saved, while Mendeley had about 92%.

Spearman correlations between citations in GS and WoS were high in this research (0.957 for Nature and 0.931 for Science). The correlations between Mendeley's user count and the citations in GS and WoS were also rather good (0.559 and o.592 for WoS and GS respectively for Nature, 0.540 and 0.603 for Science). CiteULike had far weaker correlations: 0.366 with WoS and 0.396 with GS for Nature, 0.304 with WoS and 0.381 with GS for Science.

Limitations

The authors remind us that correlation isn't causation, saying they can't conclude a casual relationship based on correlations between two data sources. Therefore, it can't be determined for sure whether there is a connection between a high user count and a high number of citations. Only Nature and Science were studied, so it can very well be that the results aren't true for other journals. Also, group-saved and single-user saved references were given the same weight. The number of saved references in Mendeley and CiteULike is much smaller than in the WoS counts and therefore the results might be less reliable.

The authors speculate that user count may represent a more accurate scientific impact of articles, and take note that one can measure the impact of all sorts of resources in online reference managers, unlike in the limited bibliographic indexes.

I think it could be reference managers don't always reflect readership: one could save a reference and forget about it all together later (so many articles, so little time...). On the other hand, citation counts might suffer from the same problem, as many scientists use a "rolling citation" from other articles citing an earlier article, without actually having read the article themselves.

Priem et al. also presented lately a study about web citations and WoS citations, based on data from the seven PLoS journals, but I think I'll wait for the journal article to cover it in the blog.



Li, X., Thelwall, M., & Giustini, D. (2011). Validating online reference managers for scholarly impact measurement Scientometrics DOI: 10.1007/s11192-011-0580-x

Friday, May 20, 2011

You're just a number: introduction to the h-index

Measuring a single scientist's output has always been problematic. Why? First, in order for the statistics to be reliable, the scientist has to produce a considerable publication output and get cited. That takes time. Second, measures like research productivity, number of publications and citations don't always correlates. Measuring the output of journals and universities has been far more reliable than measuring that of one person.

Suggested by physicist Jorge Hirsch, h-index (2005) offers an attractive way of quantifying one's scientific output as a single number. The index is defined as:

“A scientist has index h if h of his or her Np papers have at least h citations each and the other (Nph) papers have ≤ h citations each” (Hirsch, 2005).

So, if a scientist published at least ten papers, which each were cited at least ten times, her h-index is ten. A zero h-index, on the other hand, says that the scientist perhaps published papers, but is yet to have an actual impact.

The h-index is attractive because it takes into account both the number of publications and the number of citations. It isn't phased by "one hit wonders", but favors a body of work that each of its components has at least a certain impact (citations).

Problems and disadvantages

Which database to use? Different databases cover different journals, conferences, etc. Web of Science, for example, has better coverage of STEM than of the humanities, which tend to publish books rather than papers. Using Google Scholar will likely inflate the h-index.

Which field are you in? Larger fields mean a larger potential for citations, resulting in a higher h-index.

You aren't a number! (Or at least, not just *one* number). Reducing scientists to a single number ignores other factors, such as their teaching skills and ability to collaborate. Can an entire career really be described as a single number?




The age factor: The older the scientist gets, the longer she had to publish and get cited. Younger scientists are at disadvantage with the h-index.

Relevance: Since the h-index doesn't decrease, it can't tell whether a scientist is still active and/or where her work is still relevant for others in her field.

Since the h-index is a single number, scientists with the same h-index can have very different numbers of papers and citations. In the following table, scientist A and scientists B have the same h-index, but scientist A has far more citations in the overall raw calculation.



Because of the h-index many problems, offering new corrections to it, or coming up with other indices altogether is the official new sport for bibliometricians. The new indices are supposed to offer a better way to make decisions about promotions and grants, but despite all the efforts, it seems that the way to the promised tenure will continue to be paved with peer evaluation.

Bornmann, L., & Daniel, H. (2007). What do we know about the h-index? Journal of the American Society for Information Science and Technology, 58 (9), 1381-1385 DOI:10.1002/asi.20609

Bornmann, L., & Daniel, H. (2008). The state of h index research. Is the h index the ideal way to measure research performance? EMBO reports, 10 (1), 2-6 DOI: 10.1038/embor.2008.233

Hirsch, J. (2005). An index to quantify an individual's scientific research output Proceedings of the National Academy of Sciences, 102 (46), 16569-16572 DOI: 10.1073/pnas.0507655102

ResearchBlogging.org

Tuesday, October 5, 2010

When is webometrics most useful?

Like many terms in Information Science (including 'Information Science' itself) the term 'webometrics' is pretty vague. Björneborn and Ingwersen (2004) defined webometrics as "the study of the quantitative aspects of the construction and use of information resources, structures and technologies on the Web drawing on bibliometric and informetric approaches." I guess this definition will have to do for the time being.

Thelwall*, Klitkou, Verbeek, Stuart and Vincent (2010) set out to find in which fields webometrics is most effective. The result is quite a long paper, that I'm going to be very general about its conclusions. As expected, webometrics doesn't have the same effectiveness in every field. It is at its best with emerging and/or "hot" fields. That is because web publication is easier and faster than publication in traditional scientific outlets, and researchers can publish ongoing results with little delay.

In some disciplines plenty of their products aren't regularly published in journals (social sciences, humanities, applied fields, etc.) and therefore aren't as well-covered by bibliometrical databases as disciplines with a journal-publishing culture. Bibliometrics is also bound to have a poor coverage of multidisciplinary fields, because their outputs are published in various outlets and are often cited in different manners.

In general, webometrics analysis gives better results in fields with standards and/or norms for web publishing, but the results might not be reliable in fields where a small number of research groups and their projects (databases, web portals and so on) have a disproportional web presence.

Collaborations are often better caught in webometric analysis, since not all collaborative works have "official" outputs.

Webometrics works better for smaller fields. It's harder to get a complete picture of large fields with current methods.

Last but not least: webometric analysis is usually faster and cheaper than bibliometric one.

Thelwall and his colleagues concluded that "whilst webometrics is still inferior to bibliometrics for most purposes it seems that it has advantages for some types of field, particularly new, small fields, and can deliver policy-relevant (process) indicators to promote effective collaboration and communication," (in short: use with caution).


Appropriate disclosure: Prof. Thelwall is one of my dissertation advisors . My favorite from his long list of achievements is that he managed to publish a serious research paper including a YouTube cat video.


Thelwall, M., Klitkou, A., Verbeek, A., Stuart, D., & Vincent, C. (2010). Policy-relevant Webometrics for individual scientific fields Journal of the American Society for Information Science and Technology, 61 (7), 1464-1475 DOI: 10.1002/asi.21345

ResearchBlogging.org