2012-08-25

Impact of open data

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
Rufus Pollock from the Open Knowledge Foundation argues that “open data is a means to an end, not an end in itself” [1]. Open data alone has no impact, as its impact is triggered by its use. Thus, no impact is guaranteed by the intrinsic properties of open data.
Open data discourse contains a vision that promises a better society in the offing. It is a vision that stems from the belief in transformative effects of open data principles and information technologies that are entrusted to deliver this vision. However, this vision will not be put into practice by releasing open data. Its the use of open data that puts the transformation into motion.
Rhetoric of open data advocates emphasizes the positive side of open access to public sector data. Moreover, it is often presented as an asymptomatic and strictly apolitical issue. However, it would be short-sighted to assume it is a neutral, technological change. We need to admit that there are both positive and negative impacts of open data, bringing both benefits and repercussions.
Distinguishing between the target of open data impacts, a rough categorization can be drawn classifying impacts either as internal, if they affect data producers, or as external, if they influence others.

Internal impact

Internal impact, which affects the producers of public sector data, is based largely on data about the public sector. The data describing the public sector is a record of its activity that may be used and scrutinized to improve the workings of the public sector. An open and better performing public sector is among the key objectives of the open data movement. Ultimately, open data paves the way to an open and more efficient government.
Open data disrupts existing workflows that are established in the public sector. It subjects the public sector to a greater transparency, which enables to held civil servants accountable, and establishes conditions under which the public sector may function in a more efficient way.

External impact

External impact of open data affects the demand side of open data. It results chiefly from availability of data about the environment governed by the public sector bodies releasing the data.
A recognized issue with the open data movement is that it lacks focus on the demand side of data. It suffers from unrealistic expectations brought about with the pervasive tendency to pay attention solely to the supply side, which is coupled with a lack of consideration of how the data would be used after its release [2, p. 1]. The public sector should abandon this ill-considered model and instead adopt a user-centric model for data disclosure.
Close attention to the demand side is needed because the power of open data is not in itself, it resides in the ways it can empower people that use the data. Open data empowers citizens to make better decisions. For example, access to crime data may assist city dwellers in finding the safest route home. Information about wheelchair access to public transportation may help persons with reduced mobility to arrange their city transport better. The effects of open data that impact users of data are covered in the following sections. Among the effects that are discussed is the phenomenon of disintermediation that allows users of data to by-pass intermediaries and the ways in which open data enables citizens to participate in public affairs. Influences of open data on two specific domains are considered. The availability of public sector data is a new potential for the economy. For journalism open data brings about a change that makes it become more data-driven.

References

  1. POLLOCK, Rufus. Open data: a means to an end, not an end in itself [online]. September 15th, 2011 [cit. 2012-04-06]. Available from WWW: http://blog.okfn.org/2011/09/15/open-data-a-means-to-an-end-not-an-end-in-itself/
  2. MCCLEAN, Tom. Not with a bang but with a whimper: the politics of accountability and open data in the UK. In HAGOPIAN, Frances; HONIG, Bonnie (eds.). American Political Science Association Annual Meeting Papers, Seattle, Washington, 1 — 4 September 2011 [online]. Washington (DC): American Political Science Association, 2011 [cit. 2012-04-19]. Also available from WWW: http://ssrn.com/abstract=1899790

2012-08-24

Linked open data in the public sector

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
Having reviewed the theoretical foundations for technical openness and data quality of linked data, this section turns to the ways in which linked open data is used in practice in the public sector. Contrary to the popular belief, linked open data is not any more confined to the research institutes producing pilots and prototypes. It is used in practice, and the public sector is one of the central areas in which linked data is being adopted.
To find out about the role of public sector data in the ever-increasing web of data, the Linked Open Data Cloud diagram may be consulted. This diagram depicts the connections between the existing linked data sources that are published under the terms of an open licence. Progressive changes made to this diagram over time illustrate the growth of the web of data that now contains more than a billion triples. The cloud is partitioned in broad subject categories that include a category for “government”. According to the State of the LOD Cloud [1] survey from September 2011 the datasets in this category represented 42.09 % of triples in the cloud. However, these datasets accounted only for 3.84 % of outbound links to external datasets.
The Linked Open Data Cloud features datasets from the public sector of a number of countries. The U.S. is represented by their pioneering Data.gov project started by the Obama administration in May 2009. In the United Kingdom, the adoption of linked open data in the public sector was kick-started by research projects, such as AKTivePSI [2]  at the University of Southampton. The research activity quickly developed into an official part of work of the public sector and gave rise to Data.gov.uk, one of the most comprehensive and progressive government data catalogues to-date. Aside from the other countries, initial experiments with linked open data for the data produced in the public sector are also conducted in the Czech Republic by an un-official initiative OpenData.cz.
The thriving growth of linked open data activities in the public sector pointed to a need for coordination and development of standards and best practices. The W3C has taken the lead and established the Government Linked Data Working Group to help guide the adoption of linked open data in the public sector. The group is scheduled to run until 2013, but it already published several recommendations, such as the Cookbook for open government linked data [3].

References

  1. BIZER, Chris; JENTZSCH, Anja; CYGANIAK, Richard. State of the LOD Cloud [online]. Version 0.3. September 19th, 2011 [cit. 2012-04-11]. Available from WWW: http://www4.wiwiss.fu-berlin.de/lodcloud/state/
  2. ALANI, Harith; CHANDLER, Peter; HALL, Wendy; O’HARA, Kieron; SHADBOLT, Nigel; SZOMSZOR, Martin. Building a pragmatic semantic web. IEEE Intelligent Systems. May—June 2008, vol. 23, iss. 3, p. 61 — 68. Also available from WWW: http://eprints.soton.ac.uk/265787/1/alani-IEEEIS08.pdf. ISSN 1541-1672. DOI 10.1109/MIS.2008.42.
  3. HYLAND, Bernardette; TERRAZAS, Boris Villazón; CAPADISLI, Sarven. Cookbook for open government linked data [online]. Last modified on February 20th, 2012 [cit. 2012-04-11]. Available from WWW: http://www.w3.org/2011/gld/wiki/Linked_Data_Cookbook

2012-08-23

Linked data: quality

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
Data quality is not inherent in technologies but it is a result of the way technologies are used. Apart from the strict limitations of semantic web technologies and linked data principles enforced by peer pressure, there is a body of knowledge about linked data captured in informal design patterns and best practices, that is embodied in resources like Linked data patterns [1] or Cookbook for open government linked data [2]. Among the other aspects these recommendations deal with they propose ways how linked data should be used to achieve the best data quality.

Content

The content facet of open data quality metrics tracks if the content of data is primary, complete, timely, and delivered intact.

Primariness

A key principle of linked data is to ensure access to raw data. Linked data URIs are required to dereference also to raw, machine-readable data, such as RDF in XML. Besides dereferencing, linked data may implement interfaces for access raw data, such as SPARQL endpoints.

Completeness

A common way to arrange for the access to complete data is to provide data dumps exported from a database or a triple store in the back-end. In this way, users are allowed to work with the data as a whole.
RDF offers an inclusive way for representing data of varying degree of structure and granularity. Depending on the modelling style, RDF can capture both highly-structured data and unstructured free-text. Linked data improves this inclusiveness by enabling to link to non-RDF content.
Linked data offers a means for materialization of the types of data that are, for the most part, out of the scope of the other approaches to data representation. For example, it may include explicit relationships between the described resources. From this perspective, linked data may be seen as a more complete representation of a particular phenomenon.

Timeliness

Even though timely release of data is rather a matter of policy and human resources, technologies employed for that task can make it easier. In particular with highly dynamic data that goes through frequent changes it is important to have a flexible update mechanism at hand. Updates of linked data may be automated with SPARQL 1.1 Update that offers a very expressive method for patching data.
Timeliness is crucial in two areas that are gaining prominence: streaming sensor data and user-generated content. Research on the technological solutions for these areas is in its infancy [3]. However, there already are experiments with streaming linked data or real-time extraction from user-generated content, such as DBPedia Live that captures updates in Wikipedia in a near real-time.

Integrity

The stack of the semantic web technologies, which linked data builds on, includes both digital signature and encryption as a part of the so-called Semantic Web Layer cake. For ensuring the content of data is not tampered with during its transmission secure HTTPS connections should be employed. An example of semantic web technology that builds on digital signatures is WebID, that may be used to authenticate data publishers.

Usability

Usability may be perceived as the weakest point of linked data. In most cases, raw, disintermediated linked data is not intended for direct consumption. This is the result of the separation of concerns that linked data employs. For example, consider working with a SPARQL endpoint that, even though it is a powerful way of interacting with data for applications, may be baffling for the regular users. Linked data should be rather mediated through end-user interfaces of web applications, that present the data in a more usable and visually-appealing manner. However, there are still aspects in which raw linked data excels when compared to other types of data.

Presentation

Intelligible presentation of linked data should be arranged for by the implementation of mechanisms for dereferecing URIs, which should be able to serve a human-readable resource representation, such as in HTML. However, representations of linked data resources are usually generated into generic templates in an automated fashion, which impedes custom adaptation of representations for different resource types.

Clarity

RDF has a well-defined way how to convey semantics through the use of RDF vocabularies and ontologies, the workings of which are described in the previous blog post about RDF. RDF vocalabularies and ontologies make thorough data modelling feasible, which increases the fidelity and clarity of the way representations of RDF resources are modelled.

Documentation

Linked data is self-describing data. Since the “consumers of Linked Data do not have the luxury of talking to a database administrator who could help them understand a schema” [2], all the information necessary to interpret the data, including RDF vocabularies and ontologies used by the data, should be stored on the Web and should be possible to retrieve via the mechanism of dereferencing by issuing HTTP GET requests and recursive following of links.
While the representations of resources should be self-documenting, there is no such requirement on the linked data URIs. URIs may be opaque since “the Web is designed so that agents communicate resource information state through representations, not identifiers” [4].

References

  1. DODDS, Leigh; DAVIS, Ian. Linked data patterns [online]. Last changed 2011-08-19 [cit. 2011-11-05]. Available from WWW: http://patterns.dataincubator.org
  2. HYLAND, Bernardette; TERRAZAS, Boris Villazón; CAPADISLI, Sarven. Cookbook for open government linked data [online]. Last modified on February 20th, 2012 [cit. 2012-04-11]. Available from WWW: http://www.w3.org/2011/gld/wiki/Linked_Data_Cookbook
  3. SEQUEDA, Juan F.; CORCHO, Oscar. Linked stream data: a position paper. In TAYLOR, Kerri; AYYAGARI, Arun; DE ROURE, David (eds.). Proceedings of the 2nd International Workshop on Semantic Sensor Networks, collocated with the 8th International Semantic Web Conference, Washington DC, USA, October 26th, 2009. Aachen: RWTH Aachen University, 2009, p. 148 — 157. CEUR workshop proceedings, vol. 552. Also available from WWW: http://oa.upm.es/5442/1/INVE_MEM_2009_64353.pdf. ISSN 1613-0073.
  4. JACOBS, Ian; WALSH, Norman (eds.). Architecture of the World Wide Web, volume 1 [online]. W3C Recommendation. December 15th, 2004 [cit. 2012-04-20]. Available from WWW: http://www.w3.org/TR/webarch/

2012-08-22

Linked data: use

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
The flexible, application-agnostic nature of linked data makes it possible to employ it for a broad spectrum of uses. Linked data does not discriminate according to the type of use as “Linked Data principles and publishing guidelines are designed to make structured data more amenable to ad hoc consumption on the Web” [1, p. 13].
Roy Fielding wrote that “the primary mechanisms for inducing reusability within architectural styles is reduction of coupling (knowledge of identity) between components and constraining the generality of component interfaces” [2, p. 35]. Fielding’s REST, covered in the previous blog post about HTTP, is based on uniform interfaces between components and thus abides by this recommendation. However, a trade-off of uniform interfaces is of efficiency because such interfaces are optimized for the general case [Ibid., p. 82]. Since linked data is based on REST it also inherits this trade-off.
Linked data adopts separation of concerns and decouples content from presentation. In this way, it decouples data from upstream (producers) and downstream (consumers) interfaces enabling variability without introducing interoperability costs. Since linked data is not application-specific it may be used to power all kinds of applications.
Modelling of linked data is based on the reuse of existing models provided by RDF vocabularies and ontologies. A common approach to modelling of linked data is to mix various vocabularies and ontologies at will, cherry-picking their components to build a customized model suited for particular data.
Flexibility of the RDF data model enables to query the data and reconfigure it for a particular use. Semantic web technologies open opportunities for reuse by offering “query interfaces for applications to access public information in a non-predefined way” [3]. This is more difficult to achieve for non-RDF data formats. For example, Fadi Maali argues that “providing the data in a fixed table structure, as in CSV files, makes it harder for consumers to re-arrange the data in a way that best fits their needs” [4, p. 86].
Together, composing data models of parts of data models already known to applications and the flexibility that allows to rearrange the data model to the application model is facilitative to generic consumption. Such an advantage is particularly manifest when applications combine multiple sources of linked data The applications of this type are referred to as “meshups” since they are built on data sources that mesh with each other [5, p. 321]. Without linked data, this scenario would require manual integration effort on the application level, whereas linked data would be already integrated on the data level.
The following paragraphs provide answers on how linked data meets the concrete criteria on the use of open data.

Non-proprietary data formats

RDF is a non-proprietary data format and its specifications are open and free for anyone to inspect and implement.

Standards

Linked data builds on web standards maintained by the W3C or the Internet Engineering Task Force (IETF). For an overview of standard specifications related to linked data see Linked Data Specifications maintained by Michael Hausenblas.

Machine readability

RDF serializations covered in the previous blog post on RDF are machine-readable. Specifications of RDF serializations have well-defined conformance criteria, which facilitate the development of standard parsers and make it possible for data to be validated for conformance, such as with the W3C RDF Validation Service.
RDF data is well-structured with a high level of granularity. Users of RDF may use it as a graph that may be broken down into individual triples, which allows access to data at a very detailed level.
Linked data makes explicit, machine-readable licensing possible by linking to licences. There are several RDF vocabularies that contain properties to do that, such as the Dublin Core Terms with dcterms:rights. For a structured representation of the licences themselves Creative Commons Rights Expression Language may be employed.

Safety

RDF cannot include executable content. Serializations of RDF are textual (with the exception of the proposed Binary RDF [6]), which promotes inspection and eases safety checks. However, using RDF in adversarial environments with security problems, such as RDF injection or query sanitization, is an area in which little research is conduced.

References

  1. HOGAN, Aidan; UMBRICH, Jürgen; HARTH, Andreas; CYGANIAK, Richard; POLLERES, Axel; DECKER, Stefan. An empirical survey of linked data conformance. In Journal of Web Semantics [in print]. 2012. Also available from WWW: http://sw.deri.org/~aidanh/docs/ldstudy12.pdf. ISSN 1570-8268. DOI 10.1016/j.websem.2012.02.001.
  2. FIELDING, Roy Thomas. Architectural styles and the design of network-based software architectures. Irvine (CA), 2000. 162 p. Dissertation (PhD.). University of California, Irvine.
  3. ACAR, Suzanne; ALONSO, José M.; NOVAK, Kevin (eds.). Improving access to government through better use of the Web [online]. W3C Interest Group Note. May 12th, 2009 [cit. 2012-04-06]. Available from WWW: http://www.w3.org/TR/egov-improving/
  4. MAALI, Fadi. Getting to the five-star: from raw data to linked government data. Galway, 2011. Masters thesis (MSc.). National University of Ireland. Digital Enterprise Research Institute.
  5. OMITOLA, Tope; KOUMENIDES, Christos L.; POPOV, Igor O.; YANG, Yang; SALVADORES, Manuel; SZOMSZOR, Martin; BERNERS-LEE, Tim; GIBBINS, Nicholas; HALL, Wendy; SCHRAEFEL, Mc; SHADBOLT, Nigel. Put in your postcode, out come the data: a case study. In AROYO, Lora; ANTONIOU, Grigoris; HYVONËN, Eero; TEN TEIJE, Annette; STUCK- ENSCHMIDT, Heiner; CABRAL, Liliana; TUDORACHE, Tania (eds.). The semantic web: research and applications, 7th Extended Semantic Web Conference, Heraklion, Crete, Greece, May 30 — June 3, 2010, Proceedings, Part I. Heidelberg: Springer, 2010. Lecture notes in computer science, 6088. ISBN 978-3-642-13485-2.
  6. FERNÁNDEZ, Javier D.; MARTÍNEZ-PRIETO, Miguel A.; GUTIERREZ, Claudio; POLLERES, Axel. Binary RDF representation for publication and exchange (HDT) [online]. W3C Member Submission. March 30th, 2011 [cit. 2012-04-24]. Available from WWW: http://www.w3.org/Submission/2011/SUBM-HDT-20110330/

2012-08-21

Linked data: permanence

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
Linked data principles enforce separation of data and applications, which promotes permanence. Modelling linked data is modelling without a context of use [1, p. 11]. When designing a data model for linked data, its creators abstract away from particular uses the data may get, such as in specific applications. Such design principle results in an application-agnostic data model that is not tightly coupled with any type of use that might be intended for the data. As a result, the data supports a wide range of unintended and unforeseen uses. Given the data is decoupled from applications using it, it needs not to be changed when the implementation of interfaces mediating it changes. Moreover, the software used for publishing or consuming linked data is in most cases open source and thus needs not to be changed if a vendor providing it changes. Even if there was no support for these open source solutions, data formats used for linked data have open specificiations that may be re-implemented by anyone.
Established design patterns for linked data promote persistent URIs providing long-lasting access points [2, p. 5]. Several of the best practices for minting URIs contribute to their persistence. URIs should not be made session-specific, in which case they cannot be used for re-identifying the requested resources after the session expires. URIs should be made implementation-agnostic because if they depend on an implementation they cannot outlast it. Therefore, URIs should not be cluttered with implementation details, such as file type suffixes (e.g., .php). A technique that further decouples URIs from the way they are dereferenced is to introduce a layer of indirection by using a service such as http://purl.org to redirect URIs to URLs that serve their representations. However, ultimately the persistence of URIs is proportional to the commitment of institutions maintaining them.

References

  1. WOOD, David (ed.). Linking government data. Heidelberg: Springer, 2011. ISBN 978-1-4614-1766-8.
  2. Designing URI sets for the UK public sector: a report from the Public Sector Information Domain of the CTO Council’s Cross-Government Enterprise Architecture [online]. 2009 [cit. 2012-02-26]. Available from WWW: http://www.cabinetoffice.gov.uk/sites/default/files/resources/designing-URI-sets-uk-public-sector.pdf

2012-08-20

Linked data: accessibility

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
Linked data requires using dereferenceable HTTP URIs that serve as open access points to data. Resolution of linked data URIs may be either implemented by serving static files or by generating resource representations on the fly.
Linked data may be published in static files in one of the RDF serializations described in the previous post about RDF. This approach is used mainly for serving RDF vocabularies and ontologies, transfer of datasets for local batch processing, or for files with embedded RDF. Serving static files is easy to implement, however, their content is fixed and difficult to manipulate and update. To take advantage of the flexible nature of linked data on-demand, dynamically generated RDF representations may be served instead. One option for this approach is to use wrappers for dynamic data extraction from non-RDF data sources. For example, D2R Server allows to expose relational databases as RDF through a pre-defined mapping.
However, to reap the full benefits of RDF a triple store should be used to store the data. Triple store is a database optimized for storage and retrieval of RDF data. To publish data from a triple store SPARQL endpoints are used as the interfaces users interact with. The endpoints expose an interface defined by the SPARQL Protocol for RDF, which allows to query or manipulate data and serves the query results in XML via HTTP. In order to comply with linked data principles publishers should use front-end applications that implement dereferencing and content negotiation. A common way how to expose RDF as linked data is through lightweight SPARQL wrappers that dereference URIs to concise bounded descriptions [1] of the requested resources, the descriptions of which they retrieve via SPARQL queries. Example implementations of linked data front-ends include Pubby or Graphite.
To ease the transition to the use of linked data for web developers specification of Linked Data API was created. Linked Data API is a framework for more user-friendly APIs interacting with linked data in a way that follows the guidelines of REST and uses simple data formats, such as JSON. Among the example implementations of this framework are Puelia and Elda.

References

  1. STICKLER, Patrick. CBD: concise bounded description [online]. W3C Member Submission. June 3rd, 2004 [cit. 2012-04-23]. Available from WWW: http://www.w3.org/Submission/CBD/

2012-08-19

Linked data: discoverability

The following post is an excerpt from my thesis entitled Linked open data for public sector information.
If we define discoverability as the ability to get to a previously unknown URI from a known URI, then this ability depends on the in-bound links from known URIs to unknown URIs. In particular, it depends on the quantity of in-bound links, how likely it is that the users will follow them, and discoverability of their referring URIs.
Linked data fulfils the basic requirement of being linkable by using static and persistent URIs. Moreover, guidelines on URI construction for linked data recommend using human-readable URIs that are easier to communicate [1, p. 4]. To increase the interconnectedness of data services were developed that take into account out-bound links as well, such as the PSI BackLinking Service for the Web of Data.
Dereferencing URIs serves as a way to discover more data. Self-describing resources of linked data “promote ad hoc discovery of information” [2]. The representations of resources the users obtain by dereferencing their URIs may contain links to other resources. This allows for a “follow your nose” link traversal exploration style, recursively navigating through the Web. Since dereferencing mechanisms adhere to a standardized protocol, it enables to automate this type of data discovery, such as with crawlers. The methods to improve discovery of linked data may be categorized either as passive or active. Passive approaches consist in publishing additional data that makes the published linked data easier to find. To improve data traversal for crawlers Semantic Sitemaps listing all the data access points may be published. Several RDF vocabularies were devised for expressing access metadata that help in data discovery, such as Vocabulary of Interlinked Datasets (VoID). A common solution for keeping a record of available data is to post data description to a data catalogue, such as the Data Hub. To address this purpose, Data Catalogue Vocabulary (DCAT) was created.
Active techniques serve the purpose of notifying linked data consumers about the existence of data. A common way to spread information about data availability is to notify prospective consumers via the ping protocol, such as with web services like Ping the Semantic Web. Submission of data to search engines works in a similar way, such as with the form for notifying Sindice, a search engine for the semantic web.
Linked data also ranks well in regular search engines. For example, Martin Moore reported that in 2010 linked data resources from the BBC’s Wildlife Finder appeared high in Google search results for animal names [3].

References

  1. Designing URI sets for the UK public sector: a report from the Public Sector Information Domain of the CTO Council’s Cross-Government Enterprise Architecture [online]. 2009 [cit. 2012-02-26]. Available from WWW: http://www.cabinetoffice.gov.uk/sites/default/files/resources/designing-URI-sets-uk-public-sector.pdf
  2. MENDELSOHN, Noah. The self-describing web [online]. W3C TAG Finding. February 7th, 2009 [cit. 2012-04-11]. Available from WWW: http://www.w3.org/2001/tag/doc/selfDescribingDocuments
  3. MOORE, Martin. 10 reasons why news organizations should use ‘linked data’. Idea Lab [online]. March 16th, 2010 [cit. 2012-04-24]. Available from WWW: http://www.pbs.org/idealab/2010/03/10-reasons-why-news-organizations-should-use-linked-data073.html