About Me
- Jeremy
- Web person at the Imperial War Museum, just completed PhD about digital sustainability in museums (the original motivation for this blog was as my research diary). Posting occasionally, and usually museum tech stuff but prone to stray. I welcome comments if you want to take anything further. These are my opinions and should not be attributed to my employer or anyone else (unless they thought of them too). Twitter: @jottevanger
Friday, September 23, 2011
Europeana's Data Exchange Agreement: more than just a licensing change
For Europeana to be more than just a portal website it must be part of the data landscape, an ever-growing circle on that famous Linked Data diagram. I don’t think anyone concerned with the project could conceive of it as anything else nowadays. The “Distribute” strand of the 2011-15 Strategic Plan (PDF) pivots on the re-use of content away from the Europeana site, via the API or Linked Data channels. This is the Europeana that I am talking about sustaining, and this is what the EDEA makes possible. Without it, the offer could only be the ghost of that: a portal site. With it, it could become pervasive. It’s the difference between a drinks machine and mains water supply. Coca Cola like to slap on a label saying “Buxton” and sell water for 10,000 what it costs from the tap to those people that come to the drinks machine. Thames Water, by contrast, put it in pipes and get it everywhere – unbranded, plentiful, pervasive. The EDEA brings into focus a similar choice for content providers: do they want their water in the mains, or would they rather put it in bottles and try to find buyers on their own? Europeana has staked its future - I do mean that - on the mains supply approach.
The two most significant resources that Europeana relies upon are finance and content*. You might also add users, or you might prefer to count them on the “benefit” side of a cost/benefit equation. Crudely put, finance comes largely from the European Commission and content largely comes from cultural heritage content partners (either direct data providers or, more commonly, via aggregators). These two stakeholder groups, providing the grist to Europeana’s mill, are looking for different things, although their interests overlap greatly. To sustain a social enterprise, which is at heart concerned with generating public good rather than a net profit, requires whoever is providing the resources for it to perceive that the value thus generated is worth the cost. If this value proposition is not strong enough to persuade them to supply the necessary resources, well, it won’t be resourced and will fail.
Europeana has now laid out its value proposition to content providers, along with what it requires in return. Given the vision of the Strategic Plan's “Distribute” strand, this requirement could never really be less than the CC0 license the EDEA requires, but equally the “mains supply” model of cultural content is the foundation of the value proposition that the EDEA offers to museums, libraries and archives that sign the new agreement: you will find your content hosed across the internet, heliping you to deliver your public mandate but also carrying new traffic back to your site and opening up commercial opportunities. The word “exchange” in the title of the Agreement also highlights the other part of the value proposition: that the enrichments Europeana makes to the raw metadata it receives will also be CC0 licensed and will be given back to the providers. Note also that whilst the EDEA is not a contract with the funders, it again reflects the value proposition to them: culture on tap through openly licenced, reusable data.
So the EDEA is not just a new licensing agreement. If you like, it is the sustainability plan in super-brief form: give us your content and let us set it free, and here’s what you get in return - oh and funders, this is why we're worth the €€€. Since the Strategic Plan laid out a vision of open-licensed, distributed and open data, this was really the inescapable conclusion. If it finds sufficient favour, that vision of a mains supply of culture may have a long life, and that bubble on the Linked Data diagram may keep growing. If it doesn’t, well, it’s either back to the drinks machine or back to nothing. It's worth the gamble, I reckon.
* Oh, and political support. This is a big one, and is the source of powerful backing for the open data train to which Europeana has hitched its wagon. If Commissioner Neelie Kroes has anything to do with it this train is going to gather speed rapidly – see her speech to the OpenEurope Summit this week for starters.
Saturday, July 09, 2011
Back to the start
In particular it's interesting to note how much of the current vision genuinely was present at the very start. I think people have sometimes assumed that EDL/Europeana were really conceived as a web portal first and a hub for content reuse second, and probably for some motivations that would make the nose of a John Birchian twitch frenetically. But from its very inception it was truly about encouraging interoperability in support of reuse, of stimulating the creative economy, and squeezing more value from the continent's heritage by allowing it to flow to wherever it's needed - at least that's how I read it, and my recent reading matter reinforces that. I won't labour the point any more, I will simply quote extensively from the eContentPlus Work Plan for 2005 (which was actually a plan for work to be done/money spent from 2006, the year before EDLnet kicked off)
Cultural and scholarly institutions, including archives, libraries and museums, are developing or creating digital collections, either by digitising or by acquiring digital resources. Though often described on institutional web-sites, these resources lack visibility at European and global level, because there is insufficient interoperability between existing networks, across types of cultural organisation and collection, and across different types of content. This is aggravated by the range of different legacy systems and practices for describing cultural and scientific resources.
Effective access and re-use requires an infrastructure which can support a range of functions, including: discovery of collections and of individual items; disclosing conditions for and authenticating use; and integrating tools, such as thesauri and ontologies, to enable multilingual/multicultural access and use. Re-use (aggregating and creatively adding to this content) also requires enriched digital objects which can eventually bedelivered through these services, supporting new economic and business models and user communities.
eContentplus aims at leveraging the multilingual availability of significant assets of digital cultural, scientific and scholarly content by supporting the development of interoperable collections and objects - on which multilingual and cross-border services can be built - and by supporting solutions to facilitate exposure, discovery and retrieval of these resources. Actions should increase the opportunities and scope for accessing these resources, tackle multilingual issues, and support the emergence of enriched cultural content. There are two complementary objectives for the work in this area: promoting an enabling infrastructure in support of access and use; and stimulating content enrichment.
[quoted from eContentPlus work programme 2005 p.10]
Thursday, June 09, 2011
Hack4Europe London, by your oEmbedded reporter
oEmbed for Europeana
I took with me a few things I'd worked on already and some ideas of what I wanted to expand. One that I'd got underway involved oEmbed.
If you haven't come across it before, oEmbed is a protocol and lightweight format for accessing metadata about media. I've been playing with it recently, weighing it up against MediaRSS, and it really has its merits. The idea is that you can send a URL of a regular HTML page to an oEmbed endpoint and it will send you back all you need to know to embed the main media item that's on that page. Flickr, YouTube and various other sites offer it, and I'd been playing with it as a means of distributing media from our websites at IWM. Its main advantages are that it's lightweight, usually available as JSON (ideally with callbacks, to avoid cross-domain issues), and most importantly of all, that media from many different sites are presented in the same form. This makes it
easier to mix them up. MediaRSS is also cool, holds multiple objects (unlike oEmbed), and is quite widespread.
I've made a javascript library that lets you treat MediaRSS and oEmbed the same so you can mix media from lots of sources as generic media objects, which seemed like a good starting point for taking Europeana content (or for that matter IWM content) and contextualising it with media from elsewhere. The main thing missing was an oEmbed service for Europeana. What you have instead is an OpenSearch feed (available as JSON, but without the ability to return a specific item, and without callbacks) and a richer SRW record for individual items. This is XML only. Neither option is easily mapped to common media attributes, at least not to the casual developer, so before the hackday I knocked together a simple oEmbed service. You send it the URL of the item you like on Europeana, it sends back a JSON representation of the media object (with callback, if specified), and you're done.(Incidentally I also made a richer representation using Yahoo! Pipes, which meant that the SRW was available as JSON too.)
Using the oEmbed
With a simple way of dealing with just the most core data in the Europeana record, I was then in a position to grab it "client-side" with the magic of jQuery. I'm still in n00b status with this but getting better, so I tried a few things out.
Inline embedding
First, I put simple links to regular Europeana records onto an HTML page, gave them a class name to indicate what they were, and then used jQuery to gather these and get the oEmbed. This was used to populate a carousel (too ugly to link to). An alternative also worked fine: adding a class and "title" tag to other elements. Kind of microformatty. Putting YouTube and Flickr links on the same page then results in a carousel that mixes of all of them up.
Delicious collecting
Then I bookmarked a bunch of Europeana records into Delicious and tagged them with a common tag (in my case, europeanaRecord). I also added my own note to each bookmark so I could say why I picked it. With another basic HTML page (no server-side nonsense for this) I put jQuery to work again to:
- grab the feed for my tag as JSON
- submit each link in the feed to my oEmbed service
- add to the resulting javascript object (representing a media object) a property to hold the note I put with my bookmark
- put all of these onto another pig-ugly page*, and optionally assemble them into a carousel (click the link at the top). When you get half a dozen records or more this is worthwhile. This even uglier experiment shows the note I added in Delicious attached to the item from Europeana, on the fly in your browser.
I suppose what I was doing was test driving use-cases for two enhancements to Europeana's APIs. The broader was about the things one could do if and when there is a My Europeana API. My Europeana is the user and community part of the service, and at some point one would hope that things that people collect, annotate, tag, upload etc will be accessible through a read/write API for reuse in other contexts. Whilst waiting for a UGC API, though, I showed myself that one can use something as simple as Delicious to do the collecting and add some basic UGC to it (tags and note). The narrower enhancement would be an oEmbed service, and oddly I think it's this narrower one that came out stronger, because it's so easy to see how it can help even duffer coders like me in mixing up content from multiple sources.
I didn't
What I didn't manage to do, which I'd hoped to try, was hook up bookmarking somehow with the Mashificator, which would complete the circle quite nicely, or get round to using the enriched metadata that Europeana has now made available including good period and date terms, lots of geo data, and multilingual annotations. These would be great for turning a set of Delicious bookmarked records into a timeline, a map, a word-cloud etc. Perhaps that's next. And finally, it would be pretty trivial to create oEmbed services for various other museum APIs I know and to make mixing up their collection on your page as easy as this, with just a bit of jQuery and Delicious.
Working with Jan
Earlier in the day I spent some time working with Jan Molendijk, Europeana's Technical Director, working on some improvements to a mechanism he's built for curating search results and outputting static HTML pages. It's primarily a tool for Europeana's own staff but I think we improved the experience of assembling/curating a set, and again I got to strech my legs a little with jQUery, learning all the time. He decided to use Delicious too to hold searches, which themselves can be grouped by tags and assembled into super-sets of sets. It was a pleasure and a privilege to work with the driving force behind Europeana's technical team; who better to sit by than the guy responsible for the API?
*actually this one uses the Yahoo! Pipe coz the file using the oEmbed is a bit of a mess still but it does the same thing
Monday, October 18, 2010
Open Culture 2010 ruminations #2: Europeana, UGC and the API, plus a bit of "what are we here for?"
Some background: Jill Cousins, Europeana's Director, outlined the four objectives that drive project/service/network/dream (take your pick), which go approximately like this:
- To Aggregate – bringing everything together in one place, interoperable, rich (or rich enough) and multilingual
- To Facilitate – to encourage innovation in digital heritage, to stimulate the digital economy, to bring greater understanding between the people of Europe, to build an amazing network of partners and friends
- To Distribute – code, data, content
- To Engage – to put the content into forms that engage people, wherever they may be and however they want to use it
A key plank in the distribution strategy is the API. For engagement, an emerging social strategy includes opportunities for users to react to and create content themselves, and to channel the content to external social sites (Facebook and the like). Both of these things are too big to go into here, but I think one thing we haven't got covered properly is the overlap between the two. Channeling content to people on 3rd party sites ticks the "distribution" box but whilst that in iteself may be engaging it is not the same as being "social" or facilitating UGC there. If people have to come to our portal to react they simply won't. In other words, our content API has to be accompanied by a UGC API - read and write. Even if the "write" part does nothing more than allow favouriting and tagging it will make it possible to really engage with Europeana from Facebook etc.
What falls out of this is my answer to one of Stefan Gradmann's questions to WP3 (the technical working party). Stefan asked, do we want to recommend that work progresses on authentication/authorization mechanisms (OAuth/OpenID, Shibboleth etc) for the Danube release (July 2011)? My answer is a firm "yes". Until that is sorted out we can't have a "social" API to support Europeana's engagement objective off-site, and if such interaction not possible off-site then we're really not making the most of the "distribution" strand either.
Open Culture 2010 ruminations #1: Linked Data
Linked (Open) Data was a constant refrain at the meetings (OK, not at the WP1 meeting) and the conference, and two things struck me. Firstly, there’s still lots of emphasis on creating out-bound links and little discussion of the trickier(?) issue of acting as a hub for inbound links, which to my mind is every bit as important. Secondly, there’s a lot of worry about persuading content providers that it’s the right thing to do. Now the very fact that it was a topic of conversation probably means that there really is a challenge there, and it’s worth then taking some time to get our ducks in a row so we can lay out very clearly to providers why it is not going to bring the sky crashing down on their heads.
During a brainstorming session on Linked Data, the table I sat with paid quite a lot of attention to this latter issue of selling the idea to institutions. The problem needs teasing apart, though, because it has several strands – some of which I think have been answered already. We were posed the questions “Is your institution technically ready for Linked Data” and “Does it have a business issue with LD?”, but we wondered if it’s even relevant if the institution is technically ready: Europeana’s technical ability is the question, and it can step into the breach for individual institutions that aren't technically ready yet. With regard to the "business issue" question, one wonders whether such issues are around out-going links, or incoming links? Then, for inbound linkage, is it the actual fact of linkage, or the metadata at the end of the link that are more likely to be problematic? And what are people’s worries about outbound links?
What we resolved it down to in the end was that we expected people would be most worried about (a) their content being “purloined”, and (b) links to poor-quality outside data sources. But how new are these worries? Not new at all, is the answer, and Linked Data really does nothing to make them more likely to be realised, when you think about what we already enable. In fact, there’s a case to be made that not only does LD increase business opportunities but it might also increase organisations’ control over “their” data, and improve the quality of things that are done with it: letting go of your data means people don’t do a snatch-and-grab instead.
Ultimately, I think, Linked Data really doesn’t need a sales effort of its own. If Europeana has won people over to the idea of an API and the letting-go of metadata that it implies, then Linked Data is nothing at all to worry about. What does it add to what the API and HTML pages already do? Two things:
- A commitment to giving resources a URI (for all intents and purposes, read “stable URL”), which they should have for the HTML representation anyway. In fact, the HTML page could even be at that URI and either contain the necessary data as, say, RDFa in the HTML, or through content negotiation offer it in some purer data format (say, EDM-XML).
- Links to other data sources to say “this concept/thing is the sameAs that concept/thing”. People or machines can then optionally either say “ah, I know what you mean now”, or go to that resource to learn more. Again, links are as old as the Web and, not to labour the point, are kinda implicit in its name.
So really there’s little reason to worry, especially if the API argument has already been put to bed. However I thought it might be an idea to list some ways in which we can translate the idea of LD so it’s less scary to decision-makers.
- Remember the traditional link exchange? There’s nothing new in links, and once upon a time we used to try to arrange link exchanges like a babysitting circle or something. We desperately wanted incoming links, so where’s the reason in now saying, “we’re comfortable linking out, but don’t want people linking in to our data”?
- Linked data as SEO. Organisations go to great lengths to optimise their sites so they fare well in search engine rankings. In other words, we already encourage Google, Bing and the like spider, copy and index our entire websites in the name of making them easier to discover. Now, search is fine, but it would be still better to let people use our content in more places (that’s what the API is about), and Linked Data acts like SEO for applications that could do that: if other resources link to ours, applications will “visit”.
The other thing here is that we let search engines take our content for analysis, knowing they won’t use it for republication. We should also licence our content for complete ingestion so that applications indexing it can be as powerful as possible. - It’s already out there, take control! We let go of our content the moment we put it on the web, and we all know that doing that was not just a good thing, it’s the only right thing. But whilst the only way to use it is cut-n-paste (a) it’s not reused and seen nearly as much as it should be, and (b) it’s completely out of our control, lacking our branding and “authority”, and not feeding people back to us. Paradoxically, if we make it easier to reuse our content our way than it is to cut and paste, we can change this for the better: maintain the link with the rest of our content, keep intellectual ownership, drive people back to us. Helping reuse through linked data and APIs thus potentially gives us more control.
- Get there first. There is no doubt that if we don’t offer our own records of our things in a reusable form online then bit by bit others will do it for us, and not in the way we might like. Wikipedia/DBPedia is filling up with records of artworks great and small, and will therefore be the reference URIs for many objects.
- Your objects as context. Linked data lets us surround things/concepts with context;
So if I think fears about LD should something of a non-issue, what do I think are the more important questions we should be worrying about? Basically, it’s all about what’s at the end of the reference URI and what we can let people do with it. Again, it’s really a question as much about the API as it is about Linked Data, but it’s a question Europeana needs to bottom out. How we license the use of data we’re releasing from the bounds of our sites is going to become a hotter area of debate, I reckon, with issues like:
- Is Europeana itself technically prepared to offer its contents as resources for use in the LD web? Are we ready to offer stable URIs and, where appropriate, indicate the presence of alternative URIs for objects?
- What entities will Europeana do this for? Is it just objects (relatively simple because they are frequently unique), or is it for concepts and entities that may have URIs elsewhere?
- What’s the right licence for simple reuse?
- Does that licence apply to all data fields?
- Does it apply to all providers’ data?
- Does it apply to Europeana-generated enrichments?
- Who (if anyone) gets the attribution for the data? The provider? Aggregator? Disseminator (Europeana)?
- Do we need to add legal provisions for static downloads of datasets as opposed to dynamic, API-based use of data?
Just to expand a little on the last item, the current nature of semantic web (or SW-like) applications is that the tricky operation of linking the data in your system to that in another isn't often done on the fly: often it happens once and the results ingested and indexed. Doing a SPARQL query over datasets on opposite sides of the Atlantic is a slow business you don’t want to repeat for every transaction, and joining more sets than that is something to avoid. The implication of this is that, if a third party wanted to work with a graph that spread across Europeana and their own dataset, it might be much more practical for them to ingest the relevant part of the Europeana dataset and index and query it locally. This is in contrast to the on-the-fly usage of the metadata which I suspect most people have in mind for the API. Were we be allow data downloads we might wish to add certain conditions to what they could do with the data beyond using it for querying.
In short I think most of the issues around Linked Data and Europeana are just issues around opening the data full stop. LD adds nothing especially problematic beyond what an API throws up, and in fact it's a chance to get some payback for that because it facilitates inbound links. But we need to get our ducks in a row to show organisations that there's little to be worried about and a lot to gain from letting Europeana get on with it.
Tuesday, May 12, 2009
Do the Europeana survey for purely selfish reasons
Please use Europeana.eu, the European digital library with 4 million digital objects, and join a survey to win iPodTouch.
Saturday, April 25, 2009
Catching up with Europeana v1.0 [pt.2]
So April 2nd/3rd were the kick-off meeting for Europeana 1.0, the project to take the prototype that launched last November and develop it into a full service. There may have been glitches at the launch but at the meeting there was a tremendous feeling of optimism, sustained I suppose by the knowledge that those glitches were history, and by the strength of the vision that has matured in people's minds.
The meeting was about getting the various re-shuffled (and trimmed) work-groups organised, with their scope understood by their members and refined in some initial discussions before the proper work begins. There are tight dependencies going in all directions between the work-groups. My problem was, on reflection, a very encouraging one: it was difficult to decide which WG I should work with, since they nearly all now have some mention of APIs in their core tasks. Given that concern over APIs was the reason I got involved with Europeana, it's great to see how central a place they occupy in the plans for v1.0. Not surprising, perhaps, given the attitudes I've discovered since joining, but feeling more real now that they're boosted up the agenda. For those who worry (as I used to) that Europeana was all about a portal this shows that fear is groundless. Jill Cousins (the project's director) distilled the essence of Europeana's purpose as being an aggregator, distributor, catalyst, innovator and facilitator; the portal, whilst necessary, is but a small part of this vision.
In the end I elected to join WG3.3, which will develop the technical specs of the service, including APIs. Jill is also organising a group to work up the user requirements (to feed to WG3.3), which I'll participate in. I guess this will also help to co-ordinate all the other API-related activity, and I'm thrilled to see several great names on the list for that group, not least Fiona Romeo of the National Maritime Museum. Hi Fiona! I hope to see more from the UK museum tech community raising their hand to contribute to a project that's actually going to do something, but for now it's great to have this vote of confidence from the museum that puts many of us to shame for their attitude and their actions.
So we heard about the phasing of developments; about the "Danube" and "Rhine" releases planned for the next two years; about the flotilla of projects like EuropeanaLocal, ApeNet, Judaica, Biodiversity Heritage Library, and especially EuropeanaConnect (a monster of a project supplying some core semantic and multilingual technology, and content too); and about the sandbox environment that will in due course be opened up to developers to test out Europeana, share code and develop new ideas. Though we await more details, this last item is particularly exciting for people like me, who will have the chance to both play with the contents and perhaps contribute to the codebase of Europeana itself, whilst becoming part of a community of like-minded digi-culture heads.
Man, you know, I've got so much stuff in my notes about specific presentations and discussions but you don't want all that so here's the wrap. As you can tell I've come away feeling pretty positive about the shape it's all taking, but there are undoubtedly big challenges, in terms of achieving detailed aims in areas like semantic search and multilinguality, but also in ensuring the long-term viability of the service Europeana hopes to supply; nevertheless the plans are good and, crucially, there are big rewards even if some ambitions aren't realised.
Within the UK there are a number of large museums with great digital teams and programmes that are not yet part of Europeana. There are also, obviously, lots of smaller ones with arguably even more to gain from being in it, but they have more of a practical challenge to participation right now. But why is it that those big fish are not on board yet? Is it just too early for them, or are there major deterrents at work? I know that there are people out there, including friends of mine, who are sceptical of Europeana's chances of success and sometimes of its validity as an idea. The former is still fair enough I suppose, or at least the long-term prospects are hard to predict; the latter, though, still mystifies me. If we want cross-collection, cross-domain search - and other functionality - based on the structured content of large numbers of institutions, there's really no alternative to bringing the metadata (not the content) into one place. Google and the like are not adequate stand-ins, despite their undoubtable power and despite the future potential for enabling more passive means of aggregation by getting, say, Yahoo! to take content off the page with POSH of some sort (which certainly gets my vote, but again relies on agreed standards). Mike Ellis and Dan Zambonini, and I myself separately, have done experiments with this sort of scraping into a centralised index, turning the formal aggregation model around, and there's something in that approach, it's true. Federated search is no panacea given that it requires an API from each content holder and is inferior for a plethora of reasons. Both are good approaches in their own ways and for the right problem - as Mike often reminds us, we can do a lot with relatively little effort and needn't get fixated on delivering the perfect heavyweight system if quick and light is going to get us most of the way sooner and cheaper. But I can't help but detect some sort of submerged philosophical or attitudinal* objection to putting content into Europeana - a big, sophisticated, and (perhaps the greatest sin of all) a European service. I sense a paranoia that being part of it could somehow reduce our own control of our content or make us seem less clever by doing things we haven't done, even if we're otherwise agile clever web teams in big and influential museums. But the fact is that a single museum is by definition incapable to doing this, and if you believe in network effects, in the wisdom of crowds, in the virtues of having many answers to a question available in one place, then you need also to accept that your content and your museum should be part of that crowd, a node in that network, an answer amongst many. If your stuff is good, it will be found. Stay out of the crowd and you don't become more conspicuous, you become less so. Time will doubtless throw up other solutions to this challenge, but right now a platform for building countless cultural heritage applications on top of content from across Europe (and beyond?) looks pretty good to me. It's heavyweight, sure, but that's not innately bad.
If your heritage organisation is inside the EU but isn't part of Europeana, or if it's in it but you aren't part of the discussions that are helping to shape it, then get on board and get some influence!
Flippin' 'eck, I didn't really plan on a rant.
*is this a made up word?
Catching up with Europeana v1.0 [pt.1]
Prototyping done, the bid was assembled to develop a full-blown service, "Europeana v1.0". This bid to the European Commission was successful and just before Easter a kick-off meeting was held at the Koninklijke Bibliotheek in the Hague to initiate the project. This is actually but one of a suite a of projects under the EDLFoundation umbrella, all working in the same direction, but I guess you could say it's the one responsible for tying them together.
So how is Europeana shaping up now? Having spent three days finding out I can tell you now that I came back feeling good - and not just because I was heading straight off again on holiday. Day 1 was about travel and (obviously) a long and lovely trip to the Mauritshaus, but it ended with an hour in the company of Sjoerd Siebinga, lead developer on the project, and a session with Jill Cousins, Europeana's director. I went to see Sjoerd because I wanted to find out how Europeana's technical solution would fit with our plans at the Museum of London for a root-and-branch overhaul of our collections online delivery system. I knew that they'd be opening the source code up later this year, and I also knew that in essence what Europeana does is a superset of what we want to do, so I figured, find out if there'll be a good fit and whether there are things I could start to use or plan for now. Laughably, I thought that we might actually be able to help out by testing and developing the code further in a different environment - as if they needed me! I'll save this for another post, but in short Sjoerd took me on a tour of what they use as the core of the system (Solr) and blew me away. There are layers that they have built/will build above and below Solr that make Europeana what it is and may also prove helpful to us, but straight out of the box Solr is, quite simply, the bollocks. I've known of it for ages, but until given a tour of it didn't really grasp how it would work for us. Many, many thanks to Sjoerd for that.
Next I met with Jill for an interview for my research on digital sustainability in museums, where we dug into the roots of Europeana, its vision, key challenges, and of course sustainability (especially in terms of financial and political support). This was fascinating and revealing and added a lot to my understanding of the context of the project's birth and its fit in the historical landsacpe of EC-funded initiatives in digital/digitised cultural heritage. As a research exercise it was a test of my ability to work as an embedded researcher; one who is not just observing the processes of the project but contributing and arguing and necessarily developing opinions of his own. I really don't know how well I did in this regard - I'm not sure how often my attempts to be probing may in fact be leading, or whether my concerns with the project distort the approach I take in interviewing. Equally I don't know if this matters. A debate to expand upon another time, perhaps.
Days 2 and 3 were the kick-off meeting, and I'll put that in another post.
Tuesday, December 09, 2008
Zemanta: another channel for Europeana content?
So what is Zemanta? Well TechCrunch just wrote about the launch of its public API, and from what they say Zemanta is looks to be amongst a burgeoning sector of semantic enhancement tools - another with an API announcement this week was uClassify, and you can also look to OpenCalais, Hakia, AdaptiveBlue's BlueOrganizer and others including Yahoo!. These are tools that take in (text) content, analyise it, and identify entities within or characteristics of that text. These might be embedded into the text, or returned as recommendations, classifications, or links to related material. Sometimes we're talking about a machine-facing service, sometimes an end-user one e.g. the BlueOrganizer plugin. With Hakia and Yahoo!, these are services built on the power of their search engines. Zemanta sounds like it's squarely in this area, digesting content and returning links, images, keywords etc. from a database including (of course) Wikipedia, Amazon and Flickr. Looks like it's a plugin too.
uClassify is a little different - it learns to classify your text as you train it. I'm characterising it as a semantic enhancement technology but that may not be right in a strict sense. In any case, it will "enrich" the content you submit by putting it into categories you've assigned. That said, when I used oFaust, one of the apps built on top of its API, it took my snippet of Moby Dick and told me it was like Edgar Allen Poe, but needed work! Hmm. Whether that was down to the classifier or the training, though, I don't know.
So to go back to how Zemanta might fit in with Europeana, it's basically that we could work with them to digest our content and create relevant links to Europeana's vast (hopefully) and authoritative collection of cultural heritage content: artefacts, media, documents, people, events, and places. This is where I expect it helps to be big and standardised, as it should be easier for companies like Zemanta to work with one provider of cultural heritage content than with thousands of museums, libraries and archives.
To read more about Europeana (formerly EDL) check out my earlier posts: Europeana and EDL
Thursday, November 20, 2008
UK press on Europeana's launch
BBC online: European online library launches
Guardian: Dante to dialects: EU's online renaissance
Associated Press: European history, culture and art goes digital (well not really UK but never mind)
I'll keep editing this stuff as I find more. Plenty of concentration on the awesome content as well as the fact the site was brought to its knees by huge traffic, which I'd see as a success of sorts - best see how the traffic holds up over the next few weeks, though.
Wednesday, November 19, 2008
Europeana launch
Here are a press release and some FAQs that are apparently no longer embargoed, so fill yer boots. And by the way I've blogged plenty about this project before both as Europeana and earlier as EDL
Europeana Press Release 20/11/2008
Europeana FAQs 20/11/2008
Monday, October 13, 2008
MashLogic: worth hooking into?
I guess people are much more likely to install a plugin that can do the same for a variety of providers in preference to one that's only good for one. So if Euroepana can piggyback on something like this, its appeal could made much broader.
Tuesday, July 15, 2008
Why Gnip caught my eye: a bit more depth (just a bit)
*******************************************************************
Many thanks for taking the time to look at my brief notes, you must be a busy man so I really appreciate it. It's definitely time I tried to put some flesh on the bones because it's true, I've barely sketched the link between Gnip and my own preoccupations.
My research is looking at how museums keep their digital stuff useful; in other words, how and when we keep on trying to squeeze value out of the digital stuff we've invested in. I'm trying to put a particularly museum-y spin on it because it would be all too easy to look, for example, at general questions related to digital preservation (yawn). Hence I'm exploring the specific conditions and challenges that museums have to face, as well as the way in which they value what they hold - as a "memory institution" with a remit to preserve and to serve the public, a museum has potentially got a slightly different way of valuing what it holds, though arguably this won't really apply to digital material except in special cases (like digital art). So that's the basic thread of my research: looking at how museums can and do decide a strategy for maximising value from their digital assets, and for planning new ones.
Of course, no museum is an island (that's kind of the point of the 'net, right?) and I'm inevitably thinking a lot about the relationships between museums and other parties that might provide or use services and data to/from them - this is key to extracting value, but it's also a dependency for which we need to understand the risks. In the museum community, a lot of the talk (for a couple of decades or more, now) is about how we share our most obvious USP: our collections data. Loads of work has been done on this and yet we still seem to be a long way from the dream of a way of effectively integrating the collections of more than a few institutions. This is why I've been working with the European Digital Library/Europeana project. The reason that Gnip caught my eye was because it suggests another model for data interchange. It may be not be appropriate for the scenario of sharing collections data, and one could argue that in some ways other museum initiatives share some of its characteristics (federated search, metadata harvesters etc.), but I was interested in whether we could learn from the model of a neutral mediating agent as rather than a central pool of data. We're not short of standards but we are short of co-ordinating mechanisms that we can all trust and feel we leave us with some control over "our" data.
The actual purpose of Gnip as an exchange for social data was probably of secondary interest to me, but of interest all the same - it's just an area I don't know much about. I think that on the whole museums won't need to concerns themselves directly about how whatever it is they do will relate to Gnip: I presume that if they incorporate a third party service in their site, or perhaps have an installation of WordPress, then a lot of the mechanics may be dealt with already (or will be in due course). But concern about interoperability and data portability may well be a reason why many museums (my own included) haven't yet done an awful lot with social software - although there are some notable exceptions. If Gnip helps to address these concerns then all that will still be lacking is our imagination!
One other possibility is that museum applications could indeed work with Gnip to integrate individuals' public information with their own services - say, by drawing links between a person's list of interests or music preferences, and what's in a museum's (or a library's)collection; or by suggesting events to attend based on user location, age and interests. I don't understand Gnip well enough to know if this is plausible, though, but it's an intriguing prospect.
*******************************************************************
Thanks again to Eric for taking the time to contact me, I think it speaks well of new ventures like this (OpenCalais was another) when the key people go out out of their way to make contact with the people that are talking about them.
Wednesday, June 25, 2008
Conference ketchup
In the end, nervous anticipation gave way to the onrush of time and once I was up there in front of faces familiar and not I felt a more at ease than I would have expected. Having listened to the recordings, well, there were a lot more "ums" and "errs" than ideal, but hey, I didn't forget too many things and I kept pretty close to time, which is a big improvement on my earlier debacles.
So what was I talking about? In Leicester, I talked about Europeana. It was not meant to be an overview as such (that's not really my role), but an account of my involvement and interest, focussing on my hopes for the project and, of course, the role that APIs play in that. During Q&As and coffee breaks I had a lot of really useful feedback to my question: what is stopping many more UK museums from getting involved in the project? On the whole these revolved around the burden and mechanics of providing data, which was pretty much as I suspected. It's made me more determined to do what I can to simplify these processes, but also to ensure that the pay-off to partners is as high as it can be and as well understood as possible. Perhaps we have the furthest to go to achieve the latter.
At the Koninklijke Bibliotheek in the Hague I had an even shorter slot, which was fine by me, as part of a panel whose other members were intimidatingly illustrious. The subject of the conference was "Users expect the interoperable", and this particular session had two panels discussing interoperability in relation to archives and museums, respectively. I took part in the latter panel. I still don't know if I actually said anything, really, because I had little in the way of conclusions to offer: I just teased out some ways in which I thought "interoperability" questions pertained to APIs in a museum context. I also looked at a few examples from the world of semantic enrichment - a strange choice, perhaps, but made because there are really no proper museum APIs to compare to, and in order to show that a lack of standardisation in that area is no barrier to those APIs (Calais, Hakia, and Yahoo! Term Extractor) being useful. Simplicity gets you a long way, as does the use of existing data formats (e.g. DC or microformats). These also fit well with the other drum I was banging, the services that EDL could offer to contributors and third parties for enriching content. So, a kind of bitty talk but at least it was brief!
On Tuesday the conference wrapped up (and I do want to talk a lot more about it ASAP, because apart from anything else the first prototype was shown off and it's COOL!). I attended a hurried meeting of WP1 and Harry Verweyen presented his paper on the business model. I think he's done a great job, although this is so far outside my area of comptence I scarcely dare comment. He'd also done a lot of work integrating some of my suggestions into the plan, and it became still clearer to me how much of this hangs off the success of the semantic web tech part of the project.
Both conferences were really rewarding in their own ways and I'll try to offer some proper notes from them as soon as I find my feet again.
Friday, June 20, 2008
Hakia, semantic enrichment, and EDL
Semantic enhancement (as well as data validation) is a service that I've mooted as a possibility for Europeana. With a specialist and very authoritative data set, it could appeal to those needing to enrich cultural heritage content. There may not be a lot of money in it, but as Harry Verwayen pointed out to me, that's not necessarily the only benefit to the service provider (or I doubt Reuters would be in this game). Building traffic around the site and strengthening the brand is a benefit. Similarly, increasing the use of the ontology/thesauri used by EDL increases its influence.
Harry is putting together a presentation for next week's WG1 meeting after the EDL plenary, and he's generously put my name on the front too although I have little to contribute beyond these slightly flaky suggestions. We'll be throwing these ideas into the mix in a discussion of business models for EDL when it goes live.
Thursday, May 29, 2008
OT: Leuven photos
Created with Admarket's flickrSLiDR.
Thursday, May 22, 2008
Leuven that aside...
Anyway, to the point: bloody hell, railways! Well there are other points, but just wanted to get that off my chest. To Beligium by Eurostar is lovely, unless there's a strike. Eurostar to their credit did all they could to help, moving me onto a later train in the not-quite-certainty that it would get past Lille. They also told my hotel I'd be late. I got there, well, the right side of dawn but the wrong side of midnight. And travelling back I had the more day-to-day joy of British trains being titsup. Not forgetting Junior puking on the journey to the train before all of that. I must have offended St Christopher or something.
But the travel trauma was all worthwhile. The meeting itself was really good and I would love to return to Leuven, it's a beautiful and tranquil place. Imagine a city centre so quiet at rush hour that most of the traffic consists of parent cycling next to their infants on the way to school.
My take-homes from the meeting were many. Here are some.
- There seem still to be divergent thoughts on whether there will be a record page, although I think it's looking very likely (as seen in the maquette). This means different things for different types of institution and material. Where the original DO on the institution's website is more than an image of moderate size, the visitor will have a motivation to leave Europeana and visit the original. For films and audio this will be clearcut. For assets where the surrogate is in any case an image, they may be less likely to leave. Perhaps this depends upon the size of the largest image that Europeana will show. It does bring home, however, how vital it will be to demonstrate to content contributors that there will be superb reporting of usage of their material on that site.
- linked to the previous point, if EDL hosts only image surrogates (and occasional small derivatives of multimedia?), it keeps its costs down and traffic to institutions high, but we must keep in mind quality control of the originals hosted elsewhere - what will be acceptable standards for different formats, and how do we control this when they're actually never ingested?
- a drawback to the plan of ingesting only thumbnails is that for institutions with no existing online presence this is not very helpful. Perhaps EDL should offer a premium service, or one for limited numbers of surrogates per institution, where they can offer to hold a full size image/media asset and display this in a modified details/record page. I'm very keen on this, as a means to attract the participation of tiny museums by, essentially, offering to get them online for the first time. The current model presupposes that every partner is already online - we need a 180 degree turn on this. It's also possibly a modest source of revenue.
- EDLLocal is obviously the plan to encourage the participation of the minnows at the moment, and I need to find out more about this. Nevertheless, if it still presupposes a web destination outside of Europeana for all DOs it's not enough, IMHO. We need to see just how low we can set the barrier to participation, and how big we can make the reward. Rather than requiring a URL for each DO, it would be better to be able to take whatever a contributor can give - a DVD of images and a spreadsheet of metadata, say - and offer them
- fuller record pages on Europeana
- API access and code fragments for dropping onto a blog or whatever
- I like the idea of using a wiki for UGC related to objects. It has the advantages of being
- cheap and out of the box
- very easy to create new pages
- EDL GUIDs/DOIs usable for page names
- microformat/POSH friendly?
- familiar to many
- clearly distinct from the "authorised" content
- The new home page that Jon Purday showed looks like progress. The concept is a good step, the graphical side isn't finished but getting there
- There was plenty to read and discuss about reorganising EDL for the next phase of the project, the build of version 1.0 (first we have to launch a series of prototypes). It's too early to talk about this.
- We worked on the vision and mission, which was quite fun and threw into relief some differing ideas of what the whole thing is about. Personally I like for a vision something like: Europeana.eu: culture and heritage, connected and shared. It's short, it emphasises both the connections being made between knowledge and the sharing of this with people.
- Money is tight, realistically EDL will be relying on a subsidy of some sort for a good while to come, but there may be some good commercial opportunities. These needn't conflict with either the ownership of the source data and digital assets by contributors, nor the public service/public good ethos. The semantic graph derived from the combined dataset will belong to EDL, and this could be very marketable. I have to work up ideas here, I have OpenCalais in mind as some sort of model.
Tuesday, May 20, 2008
To Leuven (expenses paid), but who will pay for EDL?
A significant factor in the search for revenue-raising avenues is the fact that Europeana is not going to be a content owner in any significant way, but rather a broker/facilitator for accessing content owned by others. One possibility that I believe it could be worth exploring for two reasons is some form of partnership with a search provider. Yahoo! may be a bit too distracted to talk at the moment, but along with Google could be productive partners. Both sides could benefit by working on an interface and aligning their data structures, and EDL could perhaps offer quite a bit to such a partner in terms of preferential access to the semantically enriched data it will hold. This might be directly to do with searching the resources in EDL, or it might be, say, helping to clean up datasets of people and places. In exchange, maybe either some cash or technological assistance? Perhaps some of the semantic-y startups currently taking wing could also be interesting to work with, but they won't be as well resourced. Cultural heritage organisations have a lot of knowledge and context to offer here so maybe there's a business model to be had.