About Me
- Jeremy
- Web person at the Imperial War Museum, just completed PhD about digital sustainability in museums (the original motivation for this blog was as my research diary). Posting occasionally, and usually museum tech stuff but prone to stray. I welcome comments if you want to take anything further. These are my opinions and should not be attributed to my employer or anyone else (unless they thought of them too). Twitter: @jottevanger
Tuesday, April 15, 2008
A Babelfish bookmarklet
Well, there's no public API for Babelfish (at Google, Yahoo! or Altavista) as far as I can tell, so doing what I really want to do isn't going to be straightforward. Getting the text translated means receiving the results as a full HTML page, so embedding the translation alone will involve some screen-scraping. The next best thing would at least be to highlight some text and go straight to the translation, so I've made a bookmarklet for the job: trans PT_EN
If you want this, drag it to your Links bar in IE (or right-click, save to Favourites>>Links), and in Mozilla drag it to your Bookmarks toolbar (I may have remembered this wrong). By changing the language pair indicated at the end of the redirect URL you can modify this for lots of other languages (this one is pt_en i.e. Portuguese to English). Personally, I may not use it very often, I'll have to see. Let me know if it's any good for you. You'll currently need a different bookmarklet for each language pair, of course.
I hope I can make some improvements. One would be letting the user set the language pair each time - perhaps with a prompt box. Perhaps the next thing would be to pass the translated page through a Yahoo! Pipe and scrape out the translation, to drop it straight on the page.
EDIT: Oh sod it, here's a version that lets you set the language pair in a prompt: translate
EDIT AGAIN: Just seen that Google's translate page offers something pretty much the same, dammit - though not my second option with the prompt. Perhaps I should mod their code, it will be better than mine...
KML goes open
Friday, April 11, 2008
Standing back for a moment [cross-post from mymuseumoflondon.org.uk]
[cross-post from mymuseumoflondon.org.uk]
Hi, it’s the web-monkey again. Things have been pretty intense lately, due in large part to the end of the financial year and the need to wrap up all sorts of budgets.
My part in the various projects I work on ranges from major to peripheral - sometimes some serious programming, sometimes offering advice on commissioning, sometimes just doing a little tweaking ready for integrating someone else’s work. All the same I’ll flag up a couple of things I’ve been involved in lately, at least of those that have now launched, even if I didn’t do that much myself - after all, where else do we sing about some of this stuff? Too often it ends up sort of dribbling out because we’re all too busy or exhausted to make a song and dance about it. So, here we go:
- The Great Fire of London website, orientated at children of Key Stage 1 age (5-7) and their teachers. This is the result of a partnership between the Museum of London, National Portrait Gallery, The National Archives, London Metropolitan Archives, and London Fire Brigade Museum. It’s cool. Thanks to ON101 for building the game and designing the site, and our own Mariruth Leftwich for shepherding the whole thing. Also via Mariruth comes a game to complement our Digging Up the Romans learning resource.
- At last we have sort of launched “The Database of 19th Century Photographers and Allied Trades in London: 1841-1901“. This is the electronic representation of the amazing work done by David Webb in cataloguing thousands of people in that industry in Victorian times. I built the database, hmm, several years ago for another partnership we’re in, but it was never launched for reasons that even now seem obscure. Anyway, it’s now live and, though it needs an overhaul even now, it’s great to think it may at last start being useful. I want to open the data up for mash-ups….when I get some time.
- The Sainsbury Archive, a fantastic resource at Museum in Docklands, has a new site through the efforts of archivist Clare Wood
- I can’t tell you about the work I’ve been doing on republishing an archaeological reference text, because it’s not ready yet. If you can find the test URL, well, you’re very sneaky.
- Any day now we’ll see the launch of the “Family Favourites” pages on the Museum in Docklands website. Go and seek it out, there’s a fun game and an introduction to various highlights of the galleries there.
- It’s just a promo site until the exhibition itself happens, but have a look at the Jack the Ripper pages. That’s gonna be well worth a visit - get yourself some tickets!
- Geek stuff: some time ago I made a machine-friendly interface to look at the database of publications our archaeology service (MoLAS) produces. Whilst working towards the launch of http://www.museumoflondonarchaeology.org.uk/ I decided I wanted to change the architecture of the publications application, which for one thing makes it easy to drop little nuggets of info about our publications around the site, all fed from a database. The solution I went for also works for machine access by anyone, and I hope it will be just a start: we’d like to make our events available like this, and in time our collections. For the record, it’s basically REST/XML, drop us a line if you want to use it (though I imagine that it will be the collections and events that will have wider appeal - note that events already have an RSS feed, which is used on sites like docklands.co.uk).
- And check out our events programme, I’ve just uploaded the May to August programme.
Now, what have I forgotten to mention?
Of course, there’s more in the pipeline, keep your eyes on all our sites!
Thursday, April 10, 2008
Tuesday, April 08, 2008
Significant properties workshop - report
I'm not going to write up in detail all that was presented on Monday, but highlight a few things that seemed important to me, and work out a couple of thoughts/responses of my own. I haven't yet had a chance to read the papers that were sometimes referred to at the workshop (links to them here, some are huge!) so my questions may be answered there.
- JISC’s INSPECT project, run by CeRch at KCL, has set a framework for identifying and assessing the value of significant properties (SPs), and the success of their preservation; and initiated several case studies looking at SPs in the context of sets of similar file formats (still images, moving images etc) and categories of digital object (including e-learning objects and software).
- 5 broad SP “classes” (behaviour, appearance/rendering, content, context and structure) are identified by INSPECT. These don’t seem to include space to describe the “purpose” of a digital object (DO), unless this is somehow the combined result of all other SPs. But an objective such as “fun” or “communicates a KS2 concept effectively to the target audience” needs to be represented, especially for complex, service-level resources. Preserving behaviour or content but somehow failing to achieve the purpose would be to miss the point.
- Something I’m still unclear on: is it that a range of SPs are identified that can be given a value of significance for a given “medium” or format? Or is it that a set of SPs is identified for a format, and the value given according to each instance (or set of instances) submitted for presentation? In other words, it a judgement made of the significance of a property for a format/medium, or for a given preservation target?
- Once identified, SPs provide a means for measuring the success of preservation of a file format (whether the preservation activities entail migration to or from that format, or emulation of systems that support it).
- The two classes of object explored in the workshop (software and e-learning objects) are typically compound, and are much more variable than file formats. They will inherit some (potential) SPs from their components, but others (many behaviours, for example) may be implicit in the whole assemblage.
- Andrew Wilson (keynote speaker, NAA) raised the importance of authenticity. His archivists’ point of view of this concept is not identical with that in museums, or that which I'm using in my research, but it’s useful nonetheless. I have, however, already discarded it as a significant property for most museum digital resources, with the exception of the special case of DRs held as either evidence, or accessioned into collections. Archivists’ focus on informational value and “evidence” as the core measure of (and motivation for) authenticity isn’t always useful for DRs, but it is nice and clear-cut.
- The software study drew out the differences between preservation for preservation’s sake – the museum collecting approach – and preservation for use, where the outputs are the ultimate measure of success. The SPs for these scenarios differ.This paper was very interesting, and perhaps (along with the Learning Objects paper) came closest to my own concerns, but the huge variety of material under the banner of “software” clearly makes it very difficult to characterise SPs. The result is that many of those identified look more like preservation challenges than SPs in themselves. Specifically, dependencies of various sorts might count as a significant property in a “pure preservation” scenario; but in most cases they are, more likely, simply a challenge to address to maintain significant properties of other sorts, such as functionality, rendering, and the accuracy of the outputs.
- I suggested in Q&As that my reason for being interested in SPs probably differed from that of a DO-preserving project or organisation, although they have plenty in common. Andrew Wilson said that he saw the sort of preservation (sustaining value) that I was talking about as being the same as preserving in the archiving sense. I disagree, in part at least, because:
- He made the case for authenticity. This doesn’t apply when one is using SPs to help planning for good management, where we just want to make sure that we’re making best use of our resources.
- For me, SPs could prove an important approach for planning new resources, whilst for archives they are primarily for analysing what they’ve received and need to preserve (although they could in theory feed into future formats, or software purchasing decisions)
- Whilst for preservation purposes it may often be necessary to decide at a batch or format level what SPs are highly valued and hence what efforts will be invested in their maintenance, for questions of managing complex resources for active use, case-by-case decisions (based on idiosyncratic SPs?) may be the norm.
- For preservation, the “designated community” is essentially a presumptive audience, whose needs should be considered. For museums looking to maximise value from their resources, the SPs will reflect the needs of the museum itself (its business objectives and strategic aims), although ultimately various other audiences are the targets of these objectives. Perhaps there’s not so much difference here.
- Fundamental to all these differences is the fact that for archives etc, the preservation operation in which they are engaged is the core activity of the organisation. In other situations, like planning for sustainability, it is not preservation of a digital object, but its continued utility in some form (any form), i.e. the continued release of value, that counts.
These differences are largely of degree, but to me there is still a worthwhile distinction between preservation and sustainability. In a sense, preservation is the action and sustainability the continued ability to perform that action, so SPs are a way of reconciling preservation with the need for it to be sustainable. Perhaps the lack of a category that outlines the objectives, rather than the behaviour, of a digital object reflects this difference between preserving and sustaining.
Testing oneTag
Cheers, Mike!
[edit] the answer to this is it didn't work and it didn't work and it didn't work and I decided to look at the Pipe, followed that lead to Technorati and found that it hadn't updated my site's content since February, pinged it and it's now listed, but because I have very little authority (a measly 3) it won't show up with the feed that's currently in the Pipe. Bummer. Still, at least I found out that Technorati had forgotten about me!
Thursday, March 27, 2008
Chris Rusbridge on significant properties
CR has just blogged about this topic again, in the run-up to a JISC workshop that I'll unfortunately miss next week.
Wednesday, March 26, 2008
The Semantic Web now - Alex Iskold's latest great primer
Thursday, March 20, 2008
OT: awesome freestyling
Gollito and Paskowski in fullest effect. Up until recently I could imagine more ridiculous moves than were actually being pulled off by guys like this, but now, well, my imagination would be very stretched to exceed this!
Wednesday, March 19, 2008
What's new
SearchMe beta, a search engine which shows a visual results (images of the web pages) categorised (as e.g. museum, art, shopping, fishing). Quite nice. Silverlight I think. It's a bit SW (in its results clustering, for example), though how it goes about doing this I don't know, but other "semantic" search stuff has shown up lately. TextWise (small "sw", I guess) has just been reviewed by TechCrunch, which was doubtless part of the point of offering a $1m prize for suggesting uses for its technology. Hakia is another such.
Stuff I've been doing the last week or two:
- the Great Fire of London site for Key Stage 1 kids finally soft-launched.
- working on templates for the Londinium site - the bulk of my time right now
- preparing the digital republication of an out of print handbook for identifying roman pottery fabrics. I probably mentioned it before, it involved the export of Quark to PDF, the export of PDF to XML, translation via several XSLT steps and manual clean-up to TEI-Lite, and finally modification of some XSLT to display this as an HTML page. Most of this was a while ago; right now I'm getting ready for the images which will need to be embedded once all the scanned thin sections are ready.
- testing out and integrating Flash interactives with our CMS. Several are pretty much ready for launch, including two from the London Sugar and Slavery exhibition, and two games
- advising as best I can on the development of the replacement map interface for the LSS gallery
- fretting over the re-branding exercise the MoL group is engaged in. how much work is it worth doing right now to fix issues on the sites if we'll be overhauling the whole thing in the autumn?
- testing new search engine SearchMe (see above). Didn't get good results for "roman london" yet, but it's only indexed a billion pages or so....
One-stop shop for non-profits at Google
Monday, March 17, 2008
A few more Paris notes and an update
Those extra points:
Geo search
There are geographical search (and geo plus time) projects going on in eContent Plus and IST, using co-ordinates, place names, changing boundaries etc. We would hope to incorporate these (possibly post-prototype). Everything in Europeana will be public domain (development-wise) therefore the software will be there for the taking (I hope I got that right!)
"Privileged" tags
We mooted the possibility of privileged tags, i.e. those produced by certain authorised users, perhaps agreed by certain groups. Tags created by these users (most likely content contributors) would be treated differently so that we could pull out only certain items with a tag. But probably, rather than giving them some specific "privileged" status, we could achive the same thing just by identify them by contributor, user group or contributor type.
Stuff to clarify
- Licensing data model and assumptions
- Core common data
- Where is the boundary between Europeana and the contributor sites? Maquette seemed to include considerable data and the actual content displayed in-site for some types of asset e.g. images, but others might be held off-site. What are the rules?
- What needs to be added to the API to work well for libraries and archives?
Friday, March 14, 2008
Shock discovery: poor communication wastes resources
- Content maintenance. It mimics the look of our CMS pages, but the content isn't integrated with our CMS. Changes to site structure won't be reflected in the menus, nor would updated content.
- Visual maintenance. The look of this site will change (we dearly hope) and I can't change their pages
- Google. I don't know how they look upon sites that look like copies of existing sites and point at their pages. I suspect it might look like spamming and I wouldn't want to be blacklisted.
- Site stats. We can't (readily) integrate the job site's stats with ours (if we get them at all). Not a huge deal to me but a factor.
- Cost. I don't know what this will have cost, but five minutes after getting hold of an RSS feed from their site I had integrated it into our own, replicating the most important part of what they'd done. I suspect we could have done it cheaper, in short!
Thursday, March 13, 2008
Yahoo semanticises(?) business
Hmm, exciting!
[edit] see this too
The PSP is dead in the water
Tuesday, March 11, 2008
Here's that EDLNet presentation and notes
Yesterday I put the EDLNet Paris slideshow onto Slideshare, but since Scribd also let's me put up other stuff I'm putting that and the notes there too. If you can't see these coz I've screwed up the Scribd embed link or something, go here for the presentation and here for the notes
EDLNet Paris presentation:
Presentation notes:
And just in case...
The previous post with options for API input parameters is also on Scribd (UPDATED 17/3/2008)
Monday, March 10, 2008
New world speed-sailing record
Fingers crossed they might even top that at Southend today. Go Dave White!
Saturday, March 08, 2008
Public API inputs
Public API inputs and outputs
[edited 17/3/2008]
We discussed at the Paris meeting the range of parameters
that we thought that an API might need to handle to perform the sort of (public-facing)
tasks we envisaged. We didn't actually talk about output, except in regard
to the ability to specifiy return fields, but I think that this is actually
much the simpler part to work out. I've reworked our discussion, added a few
bits of my own (including the UGC bit), and split it into sections relating
to general parameters, filters for collections queries, and UGC. No doubt
lots more clarification and revision are needed and I'm pretty unclear on
some bits myself, but it's something!
Input parameters
The “profile” includes various elements defining the operation
in terms of function, languages, values and format of returned data etc. Collections
data requests will be required for some functions, and consist of various
filters. The third table relates to operations on user generated content,
including adding, editing and getting (by user or group). We may decide that
some operations are only open to specific users or categories of user; for
example, accessing UGC of some categories might only be possible for the owner
of that UGC (via their associate API key) or the owner of the collections
related to that UGC. TBC!
Query profile (data access and data addition/editing functions)
|
|
|
| A | search, compare, translate, add, update |
| A, E | DC-XML, RSS, geoRSS, CDWALite, JSON, CSV. This might instead be implicit in the target URL. |
| A | Array of field names, but a default set would perhaps include GUID, title, thumbnail, short description, owner, owner type, media. Might also provide shortcuts to preset field groups. Will vary according to target entities |
| A | Formal metadata; all data; expert and user tags; user tags only; “expert” tags only; specific user/expert/group tags. |
| A | True, false [use/don’t use thesauri etc.] |
| A, E | EN, FR etc. |
| A, E | As above. If only one is present presume the same. |
| A, E | API user key |
| A, E | For the end user. Required for accessing/modifying data attached to specific users or groups. Presumably we need to authenticate and authorise in some way, too, for some operations. |
| A, E | Perhaps multi-value, specifying rights/licensing parameters. Likely to be more complex than one field! |
Collection data filters (access only)
|
|
| Objects, people, places, subjects [if we are enabling anything more than objects] |
| A unique identifier given to every record in Europeana |
| ID for a set of entities, which may require the appropriate key, depending upon privacy settings for that set. |
| Tricycle, ww2, treaty, Anne Briggs, documentary [multi-field search]. |
| Name of object, person or place. If these use different fields, then the right one should be inferred from the target entity. Examples: photograph; sunflowers; Forlì (or Forli); Max Brod (or M Brod); Rockall |
| 14 July 1792; 19th century. “Older than 1850” might be expressed as: “- 20000000 – 1850”; “Younger than” as: “1850-2050”; uncertainty like “1850 +- 5” as a range: “1845-1855”, though this isn’t perfect |
| For returning objects, people and places |
| For returning objects, people and places. See also “geographical” below. |
| |
| Of object, principally, for documents (if this data is well expressed) |
| Good structured data, ideally (we may require an ID), but we could permit a string search across the relevant field. |
| For searching by current location |
| Museum, library, archive, A/V archive |
| Including sub-parameters for grid reference, coordinates, place name, and the location of concern (e.g. place of creation, place of publication, location of subject matter) |
| Keyword occurrence, date precision, location, location relative to user, institution type – perhaps sorting partly inferred from the fields used in search, but if these are mixed e.g. date and place plus keyword, need to sort on one before the other. |
| text, audio, image, video or more specifically PDF, WAV, MPEG etc. |
| Map, book, video perhaps. Is this data held in a structured way, and is it distinct from the media metadata? |
UGC operations (add, edit, view)
These operations will need user ID (or group ID) plus authentication and authorisation for certain operations (but not for viewing public data).
|
|
|
| A, E | For modifying tag i.e. deleting, or viewing associated items |
| A, E | For modifying or viewing |
| A, E | Perhaps multiple values, including groups, so we can look for stuff with a given tag but only when tagged by a certain set of UGC contributors |
| A, E | Content contributor vs. other user |
Friday, March 07, 2008
EDL WP3 Paris meeting
It was a successful meeting, I would say, and it was a pleasure to see a couple of faces I already knew, and meet others for the first time. On Monday (with me still reeling from a 4am start) we were taken through the results of the user testing. These were overwhelmingly positive, which needs to be taken with caution given the guided nature of the demo (especially with the online questionnaire, but also perhaps the expert users and the focus groups). All the same there were criticisms that provided something to get our teeth into, particularly around the home page and the pupose of the "who and what" tab. Search result ordering was an issue, a particularly thorny one in fact that we tackled on Tuesday as best we could. Clearly a lot of users don't really understand tagging, though they thought they liked it. Other plusses were for the timeline and map.
There was a good session with representatives of a French organisation for the blind and visually disabled after lunch (a bloomin' good lunch, in fact. Good wine, too. I love France!). Aside from HTML accessibility they talked extensively about Daisy, and it would be marvellous if some of the text content that may end up there could be daisified. No-one had heard of TEI (or DocBook) but it struck me that these formats are pretty close to what Daisy sounds like, and that there may be TEI material amongst the content we'll be aggregating, so translations to Daisy could be relatively straightforward. Anyone know?
Personalisation took us to the end of the day and we distinguished between activities done for private purposes (though perhaps with public benefits) like bookmarking with tagging, or tailoring search preferences, setting up alerts, or saving searches; and explicitly public activities like enriching content, suggesting and tagging (when not bookmarking). The question of downloads (what? how? assets or data?) and the related issue of licensing came up. I think we worked out that possibly four levels of privacy would be useful, extending the way Flickr and other sites work, with private, public, friends/family, and "share with institutions". The latter is really about saying, I will let me and my mates and the organisation whose objects I'm tagging/annotating look at the data, but not everyone. I think it's important and should be encouraged, as it lets those institutions do interesting stuff with the resulting UGC for everyone's benefit. We ran over into the next day to deal with communities (still plenty to think about there, I would say) and results display, a practical and useful discussion that touched on the fields that might be searched across and how they would be used in ranking.
Finally my bit came up. Although Fleur had suggested that I talk for maybe 15-20 minutes to kick off discussions on the API, I, feeling unsure of my ground, prepared pretty thoroughly with the result that I had material that kept me talking for an hour or more, I think, albeit with some digressions for debating what I was saying. On the whole it went down quite well, I think, but I learned a bit about what I should have added (proper, simple explanations of APIs, and more examples of how they're used) and what I should have left out (a section where for the sake of completeness I referred to the management of collection data, which is not part of the public API anyway and is outside the scope of our WP. This led to a digression that was I think still useful, but not to the topic of that moment). And then, seeding the discussion with a use case related to VLEs, we tried to figure out in more detail what functions and parameters would be needed in an API call, and what would be returned. And that, my friends, I will write up shortly for now I want my dinner. Home calls.
Thursday, March 06, 2008
MS new stuff: IE8 and super-cool image zooming in Silverlight
Also from MS, SeaDragon (see also Photosynth) in Sliverlight 2. More than a bit useful for use cultural heritage types [edit] and I should add that the video on that TechCrunch page is of a cultural heritage application - Hard Rock Cafe's memorabilia application, which was the demo shown at MIX08. They talk about the role of imaging for authentication, for bringing objects to life, and though it's obviously a business, their business is really not so far from ours (albeit for profit)
Sunday, March 02, 2008
Hey MCG Listers!
Cheers, Jeremy
Thursday, February 28, 2008
Tim Berners-Lee talks to Talis, namechecks museums (generally)
I think the Semantic Web is such a broad set of technologies and is going to doWhich is true enough, and not exactly controversial. Good to have it from the Don, though. It's all about repositories in this vision, which is fine as far as it goes but I'm not clear that it gets us all the way, though. TB-L points out that the common worry that SW means marking up HTML pages is looking at a small part of the picture and that really the bulk of the SW effort will be about databases (remember, I'm skimming!).
so many different things for different people. It is really difficult to put it
on one thing. What are the steps necessary right now for the life sciences
community to be able to use it for their data about proteins is probably
different from which steps do we need to be able to get interoperability between
repositories of library data and museum data.
My feeling at the moment is that the semantic web is building, but in good part through the efforts of those who are bypassing the "classical" technology. Inferred semantics are crucial in this phase, at least for the businesses that seem to be making some of the most interesting stuff. RDF and OWL are bit players here, although they have much more of a role in dealing with data that's already nicely structured. As TB-L points out, the steps towards full SWW do have payoffs on the way (as they must, not least in our own sector), and data integration is just such a self-rewarding step:
In fact, the gain from the Semantic Web comes much before that. So maybe we
should have written about enterprise and intra-enterprise data integration and
scientific data integration. So, I think, data integration is the name of the
game. That's happening, it's showing benefits. Public data as well; public data
is happening and it is providing the fodder for all kinds of mashups.
On the public side, light and loosely-coupled stuff will/is giving us payoffs for lower cost than going hardcore SW, and yet provides useful stepping stones on the way. Microformats, public APIs etc.: lowish cost, relatively immediate reward (potentially).
One more quote, and then I'll have to stop even skimming. This is about what to say to a CIO who want to understand what SW could do for their company, but it applies to museum people too:
"Well you should take an inventory of what you have got in the way of data and
you should think about how valuable each piece of data in the company would be
if it were available to other people across the company, or if it were available
publicly, and if it were available to your partners."And then, you should make a list of these things and tackle them in order. You should make sure you don't change the way any of your data is existing, is managed, so you don't mess up the existing systems and so on.
He then goes on to talk about the developing technological picture including SPARQL, GRDDL and the rest, and from that point on I need to read a lot more attentively....
Wednesday, February 27, 2008
The EDL API debate - Museum Computer Group thread
For those in a hurry, the quick summary of recommendations for an API is this:
- be “'open', feature-rich and based on established and agreed metadata models/standards/schemas that allow multiple sources and minimise data loss.”
- feature most of the functionality that can be accessed from the back-end
- include terms and conditions that specifically requires that UGC be flexible enough to allow any reuse with attribution
- include a key to enable differentiated access to services for different types of users
- enable the addition of “crowd-sourced” user-generated metadata
- be lightweight, using REST, XML and possibly RSS and JSON
I'm still extremely interested in any more opinions on the whys and hows of an API for EDL (or even, more generally, for any digital resource built for a museum) so please do comment or e-mail me if you have anything to add.
************************************
Summary of MCG EDL/API thread
Contributors
Jeremy Ottevanger, web developer, Museum of London
Tehmina Goskar
David Dawson, Senior Policy Adviser (Digital Futures), MLA
Mike Ellis, Solutions Architect, Eduserv
Martyn Farrows, Director, Lexara Ltd
Dr John Faithfull, Hunterian Museum, University of Glasgow
Sebastian Chan, Manager, Web Services, Powerhouse Museum
Nick Poole, Chief Executive, MDA
Terry Makewell, Technical Manager, National Museums Online Learning Project
Robert Bud, Science Museum
Matthew Cock, Head of Web, The British Museum
Douglas Tudhope, Professor, Faculty of Advanced Technology University of Glamorgan
Kate Fernie, MLA
Trevor Reynolds, Collections Registrar, English Heritage
Dylan Edgar, London Hub ICT Development Officer
Joe Cutting, consultant (ex-NMSI)
Richard Light, SGML/XML & Museum Information Consultancy (DCMI & SPECTRUM contributor, developer of MODES)
Ian Rowson, General Manager, ADLIB Information Systems
Graham Turnbull, Head of Education & Editorial, Scran
Frankie Roberto, Science Museum, London
Overview
The discussion kicked off with an introduction to EDL from JO, and a request for responses to the idea of an API for it, specifically:
- whether and why an API would be useful to them, or influence their decision on whether to contribute content to EDL
- what features might prove useful
- any examples of APIs or of their application that they think provide a model for what EDL's API could offer or enable
A second e-mail followed, offering some possible use cases for museums, libraries and archives; for strategic bodies; and for third parties.
Responses fell into three main (interconnected) strands:
- attempting to understand the role and purpose of EDL itself, and debating the value of participation
- problems relating to the practicalities of cataloguing and digitisation of collections, and the publication/aggregation of the data
- the API question
As well as providing useful ideas in respect of an API, the discussion made it clear that in the UK at least there is a need for some public relations work to be done to make the case for EDL, to explain its use for museums and to demonstrate that it will be doing something genuinely new and valuable. Barriers need to be as low as possible, and payoffs immediate and demonstrable. An alternative route to ensuring that there are contributors is coercion, so that funding is dependent upon participation, or a backdoor route wherein content aggregated for other purposes is submitted by aggregators, but ensuring institutional buy-in will be the best route to success and garner the most support. As Nick Poole (NP) himself stated:
The real question, to my mind, is whether museums perceive enough value in participating in something like the EDL to be worth the time it takes to get
involved. People have been burned in the past by services such as Cornucopia
which have tended to be relatively resource-intensive, but with little direct
payoff for individual museums - I'm not surprised people are sceptical.
EDL
Questions included, how would EDL fit in with existing EU and UK projects
such as MICHAEL, Cornucopia, and the People’s Network Discover Service. David
Dawson (DD) offered a detailed overview of its position in this network.
Cataloguing and other barriers
As John Faithfull (JF) expressed it:
I think that the current lack of killer "one stop" apps in the museum sector
is not so much due to lack of projects, technologies, or even standards, but
lack of available basic collection content for them to work with.
While supportive of APIs, he felt that it was the lack of online collection data that was the main problem. Infrastructural problems, such as access to a web server to enable automatic content harvesting in a sustainable fashion, were a big challenge. Nevertheless, he suggested that “the amount publicly available online is bizarre, bewildering and indefensible, given how technically simple the basic task has been for a long time.” Getting even flawed records out there is great for users (a point supported by Matthew Cock). Robert Bud raised some objections to this, if it just added “noise” and confusion to the internet.
NP also argued that shiny front ends tended to get financial priority over sorting out the data, but that we should get on and make the best of what we have (EDL being one means). He also felt that curators often put up resistance to getting their data online.
DD explained the planned architecture for content aggregation, which led to a discussion of software capable of acting as an OAI gateway, Trevor Reynolds pointing out that implementing an OAI gateway is not necessarily that simple. Richard Light (RL), Graham Turnbull, Ian Rowson and DD pointed to various products that do or might offer OAI servers (Modes, Scran-in-a-box, Adlib, MimsyXG and possibly others). NP indicated, too, that the solution should not be oriented at one service (EDL) or one protocol, but should be multilingual, and “the burden of responsibility has to be shifted onto the services themselves to ensure that they capture and preserve as much of the value in the underlying datasets as possible.”
Dylan Edgar pointed to the need to measure or demonstrate impact, if only in order to get funding, whilst DD reminded us that Renaissance and Designation funding, at least, came with a requirement to make metadata available to the PNDS.
An API for EDL
Mike Ellis (ME) argued that:
The notion of an API in *any* content-rich application should be moving not
only in our sphere of knowledge ("I know what an API is") but *fast* into our
sphere of requirement ("give me an API or I won't play")…
…EDL should have a
feature-rich API. A good rule of thumb for this functionality is to ask: "how
much of what can be done by back-end and developer built web systems can be done and accessed via the API?" In an ideal world it'd be 100%. If it's 0 then run
away, fast!
Applications must give us “easy, programmatic access into our data”.
Lexara’s Martyn Farrows made the case, from experience in the commercial software sector, that any API should be “'open', feature-rich and based on established and agreed metadata models/standards/schemas that allow multiple sources and minimise data loss.”
Sebastian Chan suggested that APIs may be “a *practical* alternative to the never ending (dis)agreement on 'standards'.” He suggested an API key to manage security levels and access to different services for various types of users. With regard to user generated content:
it would be prudent to have a T&C that specifically requires that UGC be
flexible enough to allow any reuse with attribution. (A CC with attribution
license may be a good option).
NP pointed out that in the cultural heritage sector the APIs of recent years have generally been one way i.e. enabling content aggregation. There is a need for evidence of the value that this returns to the content provider, in exchange for the cost of participation. He suggested that opening up the content to third parties is no different: the value is not gained directly by the content provider, and the cost of providing something adequate to all uses is probably too high. He wondered if therefore an API might be inbound as well as outbound, to allow “crowd-sourcing” of value-adding metadata creation.
JF was sceptical of the idea of working with an application housing his institution’s data, at least if this meant another obligation (providing the data):
We need stuff that makes everything easier/cheaper/faster/better rather than
having extra things to do, at extra cost.
He pointed out that the Hunterian can already do all that they wish with their own data, and doubted that any central initiative could offer much to help them add to their capacity.
Joe Cutting (JC) suggested as his main use-case the creation of exhibition displays and interactives. He indicated the problems such applications can have, such as copyright, data integrity, completeness and validity, and service level. His recommendations could be well inform an API for EDL.
In terms of technology, ME argued “lightweight every step of the way”, meaning widespread and simple technology. REST and XML (perhaps RSS too) were his preferences, rather than SOAP or JSON, which JC backed up. RL added the proviso of XML being in a community-agreed an application (for example SPECTRUM interchange format). Frankie Roberto argued for both XML and JSON, since the latter has advantages for data exchange and overcoming cross-site security issues with JavaScript.
Friday, February 22, 2008
What have I been doing?
The discussion I kicked off on the MCG list about APIs and the EDL was really stimulating, though I never managed to find the time to respond to several of the interesting responses I had. There was interesting feedback on the idea of the EDL itself, or any effort at centralising content. There were also plenty of thoughts about what make good characteristics of APIs and what their uses may be, plus suggestions of examples. I now need to use these in preparing a talk for the EDL WP3 meeting in Paris on the 3-4 March. I've been working on this quite hard, well, sort of, in between the rest, and my own ideas about APIs and indeed the Semantic Web have been evolving a little. Preparation of my survey questionnaire is another area of (out of hours) activity.
What else? I spent 3 days on preparing an out of print guide to pottery fabrics for the web. I started with a PDF of the 200 page Quark document from 10 years ago, exported XML, did some cleaning by hand, passed it through three sets of transformations to structure data and then on to TEILite, and adapted another transformation to display this as HTML. It's not perfect yet and we need to hook all the images in to it, as well as proof read etc., but it was a good proof of how one can take this content and semi-automate the creation of a nicely structured version. The plan had been to database it, and I was just going to get structured data on each fabric out but we seemed to close to what I ended up doing that I went ahead. It felt good.
I did more preparation for moving the events database out of Oracle and into SQL Server; refined some ideas with the help of the microformats list (in short, sod that, it's going to be eRDF for me or something like); met Mike Ellis to talk about the ways that Eduserve can work with museums, gab about microformats and all the rest; talked with a consultant we have working on the Port of London Archives, held at Museum in Docklands, about IT requirements, OAI-PMH and so on (followed up today with a grand assembly of all his interviewees, very interesting); worked on the London's Burning site for key stage 1 kids (more fancy pants XSLT for me, but the good bit will be a great interactive by ON 101); also work on integrating a new game, Family Favourites, soon to be launched at MiD; I got RSS feeds (and other XML sources) drawn into our site server-side (not in a public place yet, but working) and along the way rediscovered a baby REST interface I'd built in the autumn. At present this just lets you search the MoLAS publications, but events and collections are but a short way off if I get a few minutes. I advised on the procurement of a new map application for the London Sugar and Slavery gallery, although we only just made one. Still, there's money to pour away and it must be poured before the end of the financial year. Damn that financial year, it's the bane of my life right now! And I've been trying to get this blasted generic timeline project back on track. Happily today I had a meeting where we got ourselves a plan of action, and just a moment ago spoke to the designer about the next steps so now I hope there's time to get it done before, you guessed it, the end of the financial year...
Quick note: Waibel and Godby reporting on the MCN 2007 (just published in Ariadne) made some interesting observations, including: "During the public CDWA Lite Advisory Committee meeting, Inge Stein (Konrad-Zuse Zentrum für Informationstechnik, Berlin) presented on Museumdat [4], a harmonisation of CDWA (Categories for the Description of Works of Art) Lite with the CIDOC (Committee on Documentation of the International Council of Museums) Conceptual Reference Model). One of the motivations for this effort: with a small amount of changes, CDWA Lite could be used for all objects across the cultural heritage spectrum, whereas currently it is optimised exclusively for fine art. All delegates agreed that the changes proposed by the German museum community should form the basis for the next version of CDWA Lite. Monika Hagedorn-Saupe (Institute for Museums Studies, Germany), Inge Stein and the Advisory Committee agreed that a single international version of the standard would be desirable."
"The Town Hall Meeting on Intellectual Property: Museum Image Licensing – The Next Generation provoked a lively debate, with many points of view represented by both presenters and delegates, and little evidence of an emerging consensus around the business model for sustaining digital image provision: the room seemed divided between those who feel that the museum community can make the most impact in our information economy by providing open access whenever legally possible, and those who favour business models of cost recovery or even revenue generation. "
Thursday, February 21, 2008
Getting the Semantic Web to the level of the human author
Wednesday, February 20, 2008
MLA reorganisation outlined
So now we can see the plan. Well, sort of. I can't really fathom too much from the management rhetoric in here, except that they're being obliged to slim down, reshaping (for what is this, the fifth time in a decade?) and things are going to be tight. However it may all work out well, at least insofar as making it more comprehensible and transparent to outsiders. I for one have found the structure of the MLA and its agencies confusing, and this goes for the Hubs too - we may be the lead member of one such, but I still find the relationship between Hub, partners and MLA(regionalagencyhere) to be, um, enigmatic.
Hopefully the stuff about "finding new ways to share information in a digital age" is a good sign. MLA has good people on board and perhaps they'll work more tightly with frontline museum folk if their own resources are more limited. Or not - they may have less time for consultation. Time will tell.
Tuesday, February 19, 2008
Listening Post at the Science Museum
Thursday, February 14, 2008
The USA's EDL? Prob'ly not
Tuesday, February 12, 2008
Europeana demo site launched. It's cool.
Friday, February 08, 2008
TO'R and the dude from Reuters on OpenCalais and SW
Reuters CEO sees "semantic web" in its future
Thursday, February 07, 2008
I think I get Dapper now
Wednesday, February 06, 2008
OpenCalais
Monday, January 28, 2008
Nik Honeysett on planning for the "digital museum"
Thursday, January 24, 2008
OT: b3ta is the nazz
Another vote in favour of the NLP take on the Semantic Web
An interesting extra dimension to mapping
MILE - Metadata Image Library Exploitation
Monday, January 21, 2008
...and Yahoo! using user-created tags
Google as science repository
So the questions include: what form of access will there be; what tools online; what steps will this take towards semantic interoperability; and how might museums use the opportunity to make their data available, if at all?
Friday, January 18, 2008
More distributed services. Databases this time.
Thursday, January 17, 2008
Yahoo! and OpenID
OK, not much to add to this, another one on board and a very big one at that. According to RWW, this triples the number of people with and OpenID (or access to one) or will when it goes live at the end of the month.
Wednesday, January 16, 2008
Alex Iskold on "The Danger of Free"
The question we've perhaps ignored rather. It's true, we're all perhaps a bit too keen on getting stuff for free. We see stuff delivered through the screen as intangible, but that doesn't always equate to us thinking it has no value or should be free, so why is this so with web-delivered services and content? Is it just a habit, and one we can back out of if we see that it's going to mean poorer quality goods, or is it just in our nature to go for the free or to feel suspicious of the value we'll get from stuff on the web? If we have confidence that it will be trustworthy and value for money are we more likely to pay?
Wednesday, January 09, 2008
DataPortability.org gets real muscle
Bombshell: Google and Facebook Join DataPortability.org
Data portability and identity management are related to semantic web, to UGC, to interoperability and doubtless many other essential issues today, so as a user and a developer this is pretty promising news.
Museum 2.0: Setting Expectations: The Power of the Pre-Visit
A well thought out reminder of the importance of the pre-visit role of the website in contrast to the oft-emphasised post-visit role. Timely whilst we're planning the role of the digital experience in Capital City.
Thursday, December 20, 2007
The Reg: no fan of the PSP (I don't mean the Sony one)
TV licence fee 'to fund Welfare For W*nkers' The Register
I'd say not a huge fan of the Guardian either...
Monday, December 10, 2007
Cross-post: The web-monkey speaks
*************************************************
Hi. My name is Jeremy and I’m a museum web developer. There, I’ve said it. I’m a keyboard-jockey burdened (or as I see it blessed) by at least two sorts of geekdom: a late-born pleasure in ‘pooters; and a long-standing love of museums, and of the special sort of residue of our world that flocculates there (especially, it must be said, “real” things).
Seven-odd years ago I thought about setting up a museum-orientated web development company, to be called MuseioNet or some such nonsense. Thankfully it never got much beyond a cheesy name, because now I’m starting at least to get an idea of what I don’t know about the subject of building web-based resources for museums. I would have crashed and burned horribly if I’d tried to go it alone back then. I’ve been at the Museum of London about 6 years now, learning on the job. That’s pretty much inevitable anywhere, I suppose, and certainly in technology it’s a basic requirement owing to the speed of change – no matter how much you are on top of your chosen specialism today, tomorrow you’ll be slipping backwards.
Mia has already talked about her job, and although hers is more database-y and mine more webby, our roles have a fair degree of overlap so I’m not going to say much about what we do in a general sense; I want to talk instead about some current projects.
First thing to say is that roughly half of my time is spent on project-related work. Right now I’m really just wrapping up some odds and ends and catching up with some of the day-to-day stuff that’s piled up – bug fixes, data extractions, style-sheet changes. The odds and ends include a map interface for the London Sugar and Slavery website, which gives another way of exploring some of people and events that feature in the gallery of the same name. The map can be found in the gallery too. Things have been somewhat held up by our attempts to make the interface simpler and more intuitive for users, which often makes things dramatically more complex behind the scenes. I don’t know how many more
such applications we will build: we use ESRI products in-house, and their ArcIMS product powers the LSS map (and this one), but the power and flexibility of free mapping applications out there (Google and Yahoo! are amongst the most prominent) make them increasingly attractive. It would be a bit of a learning curve to learn to do all the things we want to do with these but ArcIMS is pretty complex too. We’ve already used Google Maps here and Yahoo! Maps to show our location map in context.
Another project that I dearly hope will soon be wrapped up is a very cool tool for creating quizzes and presentations, primarily for use within the context of our Learning Online site. The tool is complex and there have been problems with its development and implementation but you can see examples of what it has been used to create in the Black History section of that site. Teachers can use these interactives in a classroom situation on interactive whiteboards or regular desktop computers. The application is being developed by a Brighton company and my role has been as an advisor and in integrating it technically with our systems, as well as testing the darned thing when a fix is applied.
Mariruth Leftwich, who is overseeing the latter project, is also responsible for the recently launched poster maker, built by e-bloc to Mariruth’s brief. We can load up a bunch of images on a theme and visitors can put together a poster with them, print or submit it and finally see it in a gallery. It’s pretty cool. Again, my role was as advisor and in ensuring that it would work with our core systems and that we’d be able to live with it long-term without needing too much support. Mariruth is leading on another project I’m advising on too, which is a site for key Stage 1 kids about the Great Fire of London. This is a partnership with several other London organisations, which will fill a gaping hole of decent online resources on that subject for that age group.
The question of long-term sustainability is a key one to me, and has become so in large part because of another “project” on which I am working, namely my PhD on sustaining digital resources in museums. The museum gives me a great deal of support in this, and hopefully I am starting to return something to them as I develop as a practitioner. It’s not just me, I think that we’re all thinking a little more explicitly about the question of longevity now, bringing to the surface something that was always at least in the back of our minds. There’s a whole host of things going on in the MoL Group that tie into the question of digital sustainability: the projects I’ve mentioned and many others (not least Mia’s social software, including this blog); a review of records management; the evolving plans for IT in the Capital City Galleries that will open a couple of years hence.
Well, I’ve gone on enough for now. I should have said less about maps and more about Capital City but this is a week late already and I’d better get it online. It’ll wait till next time. Bye.
Fotowoosh breaks cover
Friday, November 30, 2007
RWW's interview with Dr. Paul Miller
Semantic Technology In Action: An Interview with Dr. Paul Miller
Monday, November 19, 2007
OpenID in HE
Friday, November 16, 2007
Brooklyn Museum and social software
Wednesday, November 14, 2007
Odds and sods 3
First, Ross's new book should be out any day now. I can't wait to read it. I just flicked through the proof in his office and know it's going to be a great and stimulating read.
We've opened a new gallery at MiD , "London, sugar and slavery", which I can't wait to see tomorrow when I'm at Museum in Docklands for the MCG meeting. My part has been to do with the ArcIMS mapping application, which isn't yet on the web but is in the gallery. Let's be honest, ArcIMS is a pain in the rear and you need a pretty good reason to justify the effort involved if you choose to use this over one of the free mapping apps, although of course they also have their learning curves and limitations. What they don't have is installation issues; OTOH you can't install them, and hence your client machines must have web access enabled. As our experiences this summer with web access on gallery machines was so dreadful we're keen to avoid this, although from past experience we know it's perfectly possible to do this safely and effectively - we just seem to be lacking the skills at present. On the subject of installation, I should say that the current version of IMS is actually pretty straightforward, perhaps disturbingly so - I think I was looking for all sorts of post-installation configuration changes to do that didn't actually need doing. But there are always complicating factors, and it's still taken me the best part of 3 days to get the thing working on our internal CMS server.
Anyway, our app uses ArcSDE, a new departure for us, and Pete's written some cool queries to make this a little more interactive than some of our previous efforts. We've got some bugs to iron out, to do with our merging and over-riding tool behaviours, but it's reasonably presentable.
Next up, Mia. Our social software torch-bearer has been working hard in all sorts of directions trying to get us off the ground with blogs, forums etc., not to mention organising our chaotic efforts with Flickr and the like. She's now got us going here: http://mymuseumoflondon.org.uk/. BIG congratulations, we're in the 21st century! Now we need to work out our management practices, encourage authors, look at how to embed and integrate this with our main sites, and see how it takes off.
I guess I should mention Jonty, but I'd rather not. If you insist you can check him out on our sites or on YouTube.
CHArt: the conference last week deserves a post of its own. For now, I'd like to give honorary mentions to J Milo Taylor, Tara Chittenden, Jon Pratty and Bridget Mackenzie, Tanya Szrajber, and Douglas Dodds, whose presentations I particularly enjoyed.
EDLNet. Did I write about this yet? I hope so. Watching over the mail list and looking at some discussion documents (so far simply lurking) I have some hope that the project will place the right emphasis on function over interface, given limited resources. Jon and Bridget talked about "Your Paintings" at CHArt, so far just a proposal but one that I would think could be designed to mesh well with EDL. I hope to talk more to Jon about this tomorrow.
Martin Bazley and Nick Poole are keen to get together with some of the people involved in IT in the London Hub so we've set up a meeting at MoL next week to see what we can draw out, initially to help them with a strategy for the SE Hub. I'm interested to see what they come out with for a strategy there. I also know that Martin wants to pursue some of the issues around stats that Dylan and I were talking about before, since he has got the job of writing a report for the London Hub on the question. I'm just a half-blind opinionated fool on the subject but if I have anything useful to offer I'll try.
Kurt Stuchell has put together a widget bringing together podcasts and blogs from/about museums worldwide. I reserve judgement on the thing itself, which I'm sure will be of use to some, possibly me included. The main point is that it's nice to see this happening in museums, full-stop. There must be lots of other imaginative ideas out there for what museum material can be widgetised. The Rijksmuseem's widget is perhaps obvious but effective nevertheless and perhaps we should do something a little similar: push object data out in an RSS feed to be consumed via a client-side JS snippet, perhaps. As I say, not that imaginative but worth a crack.
Micah Blue Smaldone. Do yourself a favour and get some. He may not be your cup of tea but you need to find out for yourself. The more I hear the more I'm ensnared. Follow far enough from the link above and you'll reach this where you can hear some spell-binding live renditions.
Thursday, November 01, 2007
What else is going on...APIs again
I'm through
Friday, October 19, 2007
Names authorities
Monday, October 08, 2007
The European Digital Library
A couple of relevant quotes:
"The Commission is promoting and co-ordinating work to build a common European 'digital library', by which we mean a common multilingual access point to Europe’s cultural heritage."
"Technology is moving fast and there are potentially many different ways of creating virtual European libraries. We should not aim at one single site or structure, but combine efforts in all the countries. What matters is to integrate access. This does not mean that the libraries or digital collections should be merged in a single database or library." [but is there one service behind it all? Or are we talking distributed search?]
"The needs of the users should be central. Developments will be demand-driven, but it is important to take a longer term and visionary view of what the user will get from the library in the way of services. Different users will have different needs and uses: One can imagine researchers wanting annotation tools; other users may wish to develop their family histories and genealogies using the materials in historical community archives." [but who is going to be able to do this development? I want to be able to point to my own sites at EDL and use its searching power and language tools to build my own applications on, which might include UGC or whatever; it would be less satisfactory to have to depend on them to build any such tools. In other words, I would want an API]
I'm still looking for clarity, then, and I've written to find out more. EDL should be fantastic, but the more open it is the more fantastic it will be, and I think that institutions will be keener to provide material if they can then hook into the back-end and really make something of it. That can only foster innovation. Fingers crossed