About Me

My photo
Web person at the Imperial War Museum, just completed PhD about digital sustainability in museums (the original motivation for this blog was as my research diary). Posting occasionally, and usually museum tech stuff but prone to stray. I welcome comments if you want to take anything further. These are my opinions and should not be attributed to my employer or anyone else (unless they thought of them too). Twitter: @jottevanger
Showing posts with label xml. Show all posts
Showing posts with label xml. Show all posts

Thursday, August 20, 2009

The great escape

Well today I don't feel like moaning. Pretty fecking remarkable, huh? Stuff went pretty well, we're close to finishing a very important stage in the Collections Online project, I unbroke some things earlier this week so I could get on with some actual work, I talked with a curator about an exciting project that's still far enough in the future that we can dream big dreams and not worry about the inevitable slap in the face that reality will give us...
On top of all that I managed to find a few minutes to do some development, which is pretty good by current standards. One thing I wanted to do was simply make a map link from an object record in our Solr index. Now, Solr URLs have their reserved characters as well as normal URL escaping. XSL, too, with which I transform the Solr output, likes escaped characters. Google Maps URLs, of the sort that you make to overlay KML on a map, well, of course they also require characters in the KML URL parameter to be escaped. The end result is a URL for a map with overlay that looks something like this:

http://maps.google.co.uk/maps?f=q&source=s_q&hl=en&geocode=&q=http:%2F%2Fwww.museumoflondon.org.uk:8080%2Fsolr%2Fselect%2F%3Fq%3Dtext:knife%2BAND%2B(start_latitude%5B-1%2BTO%2B1%5D)%26version%3D2.2%26start%3D0%26rows%3D30%26fl%3Dname,caption,start_latitude,start_longitude,site,accNum%26wt%3Dxslt%26tr%3Dkml.xsl&ie=UTF8&z=14

Ugly, huh? [BTW, once I put the new multicore index up this URL won't work]

Escape, escape, escape, and I've had plenty of fun and games in the past trying to escape stuff in XSL the way I want it without XSL then re-escaping or unescaping or otherwise ballsing up the output, so this time I thought, sod this, I'll just make a page to take in a nice simple set of parameters and redirect to the map. This makes it a whole lot easier to write the links in XSLT without worrying so much about the escape nightmare. A link like:

http://www.museumoflondon.org.uk/scripts/solrgmapredirect.asp?q=knife+AND+(start_latitude[0+TO+*])&s=0&r=30
[the "+" can be "%20" instead]

I don't know how much time I saved but I know it only took 5 minutes. It takes in a Solr query, record count and start index, escapes characters as befits GMaps KML URLs, and inserts them into a Solr query URL (including the KML transform bit, of course: wt=xsl&tr=kml.xsl, in our case). This is put into the GMaps URL and we do the response.redirect (yes, it's classic ASP). It's brittle: it will break if the GMaps URL format changes, or if the Solr URL or output format change; but hey, it's simple and works (for now).
Side benefits
It was only after making the script for these pragmatic reasons did I realise that having such a page is, of course, good for several other reasons, including:
  • it will give us stats on people following the map links
  • that same brittleness is more of a problem if I'm making links like this in lots of scripts and transformations around the site. This way I only need to point all similar links to one script and change that
  • if I decide to scrap Google Maps and use, say, OpenStreetMap, or if I want to get my KML from somewhere else, again, one script to change

I will probably add a couple of other parameters but don't want to make it heavy. Specifying the data source is one (other than Solr we can get KML out of, for example, our publications database); specifying the target service is another, so that we could use GMaps, OSM, Yahoo! and so on. Shit, anything but Streetmap (how much do you not miss having to use that piece of crap? Best thing about the last few years in mapping is the fact that you never see that anymore).

[edit 21/8/2009]

I've done some further work this morning, along the lines suggested above. It now takes in a data source and a target service parameter (though the latter only works for GMaps at present), which means I can pull in the publications KML and may start getting MOLA sites by site code too. Much more flexible now, and a single point for all map requests is going to be handy. More work to do to use more powerful aspects of pubs search.

It may seem odd to blog about a 5 minute job when I've been doing much more challenging and complex things that take months, but it's very satisfying when it works so quickly, plus my belated realisation of the useful side effects made me think it was worth talking about. Here's the salient part of the code as it now stands, for interest.

Here's a link to the new script, looking at publications data:

http://www.museumoflondon.org.uk/scripts/mapredirect.asp?r=30&q=roman&s=0&t=gmap&src=pubs

Follow that and see the GMaps URL I now longer have to write!

Sunday, March 22, 2009

Playing with SKOS

Well Mia's interest in what thesauri, word-lists etc. are out there, or could be out there, in machine-friendly form chimed nicely with mine, and it had been grating at me for ages that, for example, the NMR object type thesaurus is only available as HTML, not as a web service. There are a bunch of other thesauri in HTML form on the Collections Trust (well, MDA), English Heritage, and FISH sites, so following Mia's recent attempts to prod some of us museum tech types to action on the API front I figured I may as well have a go at turning one of them into a web service. The long and short is I haven't managed, but I have made useful steps, I think, and learnt a fair bit about Dapper, Pipes, and SKOS along the way.

I took the British Museum's material thesaurus, which is hosted by CT here. I went to Dapper and tried to get it to learn well enough to go straight to nice XML with all the different relationships having their own elements. There were too many exceptions for that and it stopped learning them after a while and I was going in circles I'd never escape, so I made a simpler Dapp (here) which just puts out the term, the linked terms, and comments. I later had to retrain it to cope with the H page but since running that page correctly once it's refused to again: it shows the results to A instead. Not to worry, add a querystring and it thinks it's a new page.

Anyway, then I had XML but still wanted to get this into nice nodes for different relationship types between terms (though wasn't really thinking about SKOS at this point. Doh!). I had high hopes for Pipes. Another doh! Because I would need to go through each item multiple times, renaming each sub-element according to its contents (e.g. broader terms all start "BT ") and trimming the string contents, I was scuppered: you can't loop operator modules, which are the ones that would allow renaming. And you can't rename by a rule, or I couldn't find how and it would probably rely on an operator module anyway. So after a lot of time wasted I thought, sod this, I know how to do this in a minute using XSLT and how important is it to have this as a web service? Fact is, it's not, or at least not in the form of a simple list - I may as well jus have a static file.

So that's what I did. It took more than a minute, though the core code scarcely did. What took longer was digging into SKOS, once it had struck me that it would be the obvious (only) format of choice. It works in a pretty straightforward way, or at least it's easy to do the basics and I didn't need to do more than that. Finding out how to represent it as RDF/XML was not so easy, coz the W3C pages don't show any - they just show TURTLE which isn't that much use to me, really. I needed a full RDF document. XML.com came up with the goods - old, but hopefully valid. So I went ahead and knocked up SKOS RDF for all the letters of the alphabet (bar X - there's nothing in the list starts with X) and merged them into one RDF file, which I hope is valid. I actually have my doubts, but I do know that with this file I can navigate around terms in a way that would be useful to me so that's good enough for me. It's here. I think it would be useful to put a web service on top of this now (perhaps Pipes can come in useful at last) so that it's really an API. Feel free! Oh, go on then, here's a first pass. Won't render as RSS and (consequently?) the "Run Pipe" screen shows nowt, but in debug mode you see results, and as e.g. JSON and PHP.

Next up there are a bunch of thesauri on those sites that I'd like to do a similar thing with, though some are going to be more fiddly. Others may be easier to dapp, but actually I reckon going to SKOS is a better bet and take it from there, as long as the content owners aren't too pissy about me playing with their stuff. Actually what would be most useful is probably to play with some of the word/term lists e.g. the RCHME Archaeological Periods List.

I could get into this.

Friday, June 06, 2008

Small API update

A couple of small advances on the API front (again, see here)


  • fixed a bug on the geo thing. For some reason an imbecilic code error wasn't breaking the script on my machine, but did on the web server. Now fixed.

  • a CDWALite-lite output for individual object records (example). There's more to add, glitches to fix, and ideally a better solution to the URL, but it's a start. Next thing is a search interface but that depends upon agreement within the Museum. A good solution may be to combine CDWALite and OpenSearch-style RSS, with the records enabling users to find the data end-point, as well as the HTML rendering. In due course I'll probably add tags to HTML record pages to point at data like this, or I may do it with some POSH.

  • the photoLondon website data now has a basic API, which I'll put on the live site next week. It returns basic person details and search parameters include: surname, forename, keyword, birth year, death year, "alive in" year, gender, country of origin, photographic occupation and non-photographic occupation. I'll work on the search result format soon, as well as the person details.

Wednesday, June 04, 2008

MoL APIs live (but very alpha)

Well there's more to do, but an alpha version of three services is now there if you want to play. All of them currently put out only XML, no JSON or raw text, GEDCOM or whatever else, but this may change and new services may be added. Have a read here. There is an events database, publications from our Archaeology Service, and a sort of geocoding tool. I've rejigged the code so that adding a different (XML) output format to these involves just writing XSLT and putting in a new value for the "mode" parameter, rather than fiddling around with C#, recompiling and all that. All feedback most welcome. But don't bug me about a collections API!

To start you off, here are links to one request from each service.
  • events API. This won't work forever as the events will expire. Uses the format that Upcoming outputs with xCal, DC and geo extensions
  • publications API
  • the geothingy converter whatsit

Saturday, May 03, 2008

PhotoLondon, genealogists and GEDCOM

Since I discovered that at least one family history website was sending users in the direction of the newly belatedly launched "Database of 19th Century Photographers and Allied Trades in London: 1841-1901", I've been thinking more and more about how we can serve this audience.

GEDCOM (for which I guess this is effectively the official homepage) seems to be the data standard of choice for interoperability in genealogy software. The latest non-XML version dates to 1995, but although its mormon keepers have been using the XML form for several years now (it was published in 2002) apparently none of the software out there in general use supports it still. How tragic is that? I guess in the museum world we're not quite the worst example of data standards paralysis! Anyway, if it had to be that which I would offer to the millions of family historians out there, so be it. It would make a lot of sense, though, to talk to that audience a bit, which I started to do on this thread. Useful feedback, not just on the worth of offering GEDCOM at all, but reminding me also of various things I'd forgotten (maybe overlooked) about that site. Copyright, sources, addresses (we have lots more structured data than is visible there), all things we could improve (given some resources).

It all ties in with an announcement this week from Stephen Brown to the MCG and MCN lists of the launch of Exhibitions of the Royal Photographic Society 1870-1915. The data in this (and an earlier site) seem so congruent with the photoLondon data that it would be lovely to explore how they might be tied together. A case for large or small semantic web, for feeds and APIs, for literally pooling data, for imaginative use of search engines or god old fashioned web content creation and management...I don't know, but perhaps we'll explore this. And now I've remembered that some of our data is more precise and structured than I had recalled, perhaps the possibilities are that much greater.

Now to get stuck into some GEDCOM. I might succumb to the XML flavour, though because frankly that 5.5 version looks like a dog.

Wednesday, July 18, 2007

The National Archives and MS

I'd wondered what to make of the news that MS are to provide a futureproofing "strategy" for TNA by giving them copies of Virtual PC 2007, apart from it being a pragmatic step, and fortunately Chris Rusbridge has done the thinking for me. I guess I agree also that the progress with Open XML can only be good, since it's pretty much irrelevant that they aren't using the same standard as Open Office - the point it it's an open standard and XML is designed to be transfomed.
I'd be interested to know how TNA plan to integrate the emulation approach with any need they may have to preserve non-Windows or non-MS Office formats.

Thursday, May 10, 2007

Forthcoming project: Slavery

A project in the works for Museum in Docklands will involve mapping slavery-related sites around London. We had a conversation yesterday about how we might accomplish the content creation for this. There are complications relating to other parties contributing content, and possibly some UGC too, but essentially we decided it was worth pursuing the idea of authoring in XML (probably TEILite, actually) since the work I've done with the format previously gives a good foundation for developing the links between sections, outputting geoRSS or KML, building indexes and glossaries etc, which would not really be an option with regular MCMS templates. Will monitor how this develops as it may turn out that the benefits don't justify the extra me-time that will be required, but there's potential there.

Extensible, reusable, Impressionable

Ben from Surface Impression yesterday came in to install the new presentation authoring tool they have built us. Things went pretty smoothly (great bloke too), we sorted out a couple of wee PHP issues relating to the old version we still run and then it all seemed to work. It's a cool app, basically we have a couple of framework SWFs that draw together a load of smaller ones for different interactions - quizzes, fill-the-gaps, matching words and pictures, video etc. - into a single "presentation" for use on the Learning Online site (or on whiteboards). Authoring is done in the same context via a collection of other SWFs and underneath it all XML is authored. We will upload the XML and collections of assets to the web server once it's all ready and pass into the framework SWF the path to the XML and hey presto. It's similar to but much more complex than other Flash stuff we've commissioned in the last couple of years and the authoring environment is cute. I'm really keen to get my teeth into the XML, actually, since I'd like to see how else we can use what is put out - a simple HTML version of a presentation should be a cinch, anyway, and it will also be easy to cut-n-paste or copy-and-edit existing presentations to create new ones.
So, we're not live with it yet but it's looking good.