Friday, May 6, 2011

Paper Reading #18: Automatic Generation of Research Trails in Web History

Comments

Reference information
Authors: Elin Rønby Pedersen, Karl Gyllstrom, Shengyin Go, Peter Jin Hong
When/Where: IUI’10, February 7–10, 2010, Hong Kong, China.

Summary
In this paper, the designers discuss the importance and motivation of constructing research trails for a variety of web searches. From an early enthographic study with ordinary people, they found that their research is generally for personal consumption, and typically occurs in a fragmented process, spread over a long time in small installments. Their research is also prone to topic sliding, where the theme may change slightly during the research process as the researcher learns more about the domain. The designers believe research trails can group together events that users perceive as belonging to the same research task and represent them as temporally ordered lists of segments.

Browsing history: one of the current tools for viewing previous web research. 

Research trails are categorized into three main entities: event, segment and topic. An event
is a page visit from a user's activity history. A segment is a temporal clustering of events. A topic is a semantic descriptor obtained from a suitable statistical/linguistic technique. Segments are formed, or clustered, when some number of minutes passes between any two consecutive events. Segments can be furthered clustered by finding average and maximum coherence, when segments share topical similarity.

A preliminary assessment of the research trail method was conducted in an end-to-end experiment involving on three users. In this experiment, we presented the segments and the research trails. For segments, we also attached associated attributes such as its time stamps, duration, topic summary, coherence values, etc. Users could explore trails and segments, and inspect the corresponding web pages. It was found that segment definition correctly modeled the natural work session for users and that trails generally contain related topics or segments.

Discussion
Initially, this reminded me of Chris's Research Field Visualizer (RFV) project for Dr. Kerne's Human-centered Computing course. It was a web application that allowed users to enter research interests and see publications related to their research field and view their published locations across the world. They can also see citation chaining between these papers. RFV seems to be a very specific form of this paper's research trail, as it is tailored for college students interested in research and forms its trails between publications based on citations and keywords, rather than between web searches and their topical coherence.

In any case, this is a very interesting and useful project. The study was rather lacking and I would have liked to preferred to hear what the users in the study thought of the idea, but I understand that the project is a work in progress. This would be useful for pretty much anyone who uses the internet to look up answers to their questions. I often find myself getting lost if my research trail becomes long. The descriptions of the four properties of early research reminded me a lot of people's low attention spans and multitasking.

No comments:

Post a Comment