Monday, 3 April 2023

Is nine million enough to fund Trove ?

 

Trove has recently announced new funding from the Australian government.

The funding works out at around $9 million Australian a year, which sounds good, but is possibly a bit on the mean side.

So why do I say that?

Trove is essentially a digital repository, which means it consists of a database containing all the metadata that is searchable and a data store that contains all the digital objects.

While it makes sense to use good quality hardware to run Trove, the hardware required is standard off the shelf stuff and not particularly exotic – commodity servers will do nicely.

The database needs to be backed up and have measures in place to ensure resilience, the object store less so as it is usually only added to. All that’s required is a periodic incremental backup in case some space junk lands in the car park, and measures to guard against disk failure.

Again, all very standard. While you need some resilience, we’re not talking about the measures the big banks deploy to ensure 24h availability of their online banking solutions.

So, the hardware part is not expensive, nor are the maintenance requirements. Electricity costs for running the servers and keeping them cool may be a constraint, but data centres usually have negotiated contracts for the supply of power, so the ongoing costs are predictable.

Likewise, the costs of periodic hardware expansion and replacement should be reasonably predictable and relatively easy to budget for.

Then there’s the costs of digitisation itself. Again, commodity hardware is now good enough for most purposes. Film scanners, microfilm scanners basically consist of standard digital camera in a housing that allows you to advance and photograph each frame.

It some cases it has to be a manual process because of the poor quality of the original film, sometimes it can be almost automated.

Scanning old fragile materials such as old newspapers, bound nineteenth century periodicals, etc is more fraught and needs both specialist skills and equipment, but I suspect that the bulk of digitisation work is the scanning of previously microfilmed material.

So, nine million should be able to cover the costs of both maintaining Trove and financing the ongoing digitisation programme as far as hardware goes.

But there’s the human factor.

Recruiting and retaining a team of computing technicians and engineers in Canberra is not cheap. I know this I’ve been there.

There's continual demand for good people and given they are all public servants or contractors, it's quite easy for people to change jobs, and of course some departments can pay more than others.

The last time I had to deal with such things was seven years ago, and the figures quoted are based on 2016 costs. Given that over the last few years wages growth has been fairly low, my costs are probably not too far out.

Once you’ve added in the costs of superannuation, long service leave, payroll tax, plus some contingency funding for sick leave, parental leave, maternity cover and the rest, a decent, competent computer technician who spends her days swapping dead disks, checking hardware status, dealing with fan failures and the like will cost around $80,000. A software engineer, between $100,000 and $120k. A service manager, at least $150k – basically your humans will cost you around a million to keep things working.

I have no idea what a competent digitisation technician costs, but I would be surprised if it was much less than a computer technician, and of course you have some more senior digitisation staff to do quality control.

I don’t know how the digitisation team is structured, but I would guess it would cost around $750k to $1million a year, meaning that overall, your human resources costs are around $2million per annum, meaning that your actual operating budget is around $7million.

That is of course in Australian dollars, and you need to maintain some wiggle room given that most serious infrastructure has a price that is tied to the US dollar price, and so even though you are paying in Australian dollars, you have to allow for drops in the value of our dollar against the greenback.

I don’t know the details of Trove’s hardware costs or data centre costs so I can’t guesstimate their annual costs. I’m guessing that the hardware consists of a few racks of servers and disk in some anonymous government data centre in Fyshwick or Hume.

So, is seven million enough?

Possibly.

The hardware and infrastructure running costs are essentially fixed costs - while its possible to lengthen replacement cycles you can only sensibly go so far, which means that your only way of reducing costs is paying people less (or paying fewer people). In Canberra you cannot really pay less than the going rate, meaning that if nine million is not enough, the headcount needs to be reduces, with consquent impacts on service and new initiatives.

I don’t know enough to say for sure. Past experience makes me feel the budget is a little tight, but not impossibly so.

However, the really good thing is that it has a recurrent and defined budget allocation. That can only be a good thing…

 

Wednesday, 8 March 2023

USB DVD's

 I've finally got off my arse with respect to retro photography.

And then I had a massive panic - what to do if the photolab wanted to deliver the scanned negatives on DVD.

We don't have any DVD drives and I wondered if we were about to have a floppy disk moment and discover we couldn't get the images read.

Part of the reason I was panicking was that we've been here before with DVD's when J needed a copy of her joint scans to take to a surgeon in Canberra who used a different X-Ray archiving service. 

(In the end the data storage people arranged to get copies printed on plastic film which solved that problem.)

However a quick search of various online market places (ebay, amazon) revealed that external USB2 and USB3 DVD drives were still available and not outrageously priced ...

Floppy disks

 



There was an article in Wired a few days ago, about how the floppy disk was still with us, living on in various embedded systems as a way of delivering software updates to elderly 747's, computer based stitching machines, and various elderly medical instruments that have trickled down to less well funded hospitals in poorer countries.

And culturally it is with us as well as a save icon in a whole gamut of applications, despite computers not having floppy drives for 20 plus years.

Of course the 3.5" 1.44MB floppy was really the end of the line* - stored in a rigid plastic jacket it was a vast improvement on the single and double sided 5.25" and 8" floppies that had come before - these really did bend, and sometimes needed support rings adding to the centre hole to avoid them being mangled by hungry drives or people pulling them out or putting them in in an overly enthusiastic manner.

But the rigid 3.5" unit was an improvement - relatively robust, and in the days before dropbox, students used to have their work on their own floppies which they (hopefully) carried about in a clean rigid plastic box designed for the purpose.

Sometimes they just used a ziploc bag, and sometimes they just stuffed it in the bottom of a backpack. It always amazed me quite how tough these floppies were and how with the aid of some canned air to blow off the crud you could usually get them to read and recover data from them.

But the great thing about the IBM format 3.5" floppy it was a near universal standard for data interchange. Every computer had a floppy drive and the disks were cheap enough to be eminently affordable.

Sometime just after the Soviet Union came apart I had an amazing demonstration on the near universality of the 3.5" floppy disk and its data format.

Among other things I was running a file and data conversion/recovery service for a university, and I remember a researcher from the Ukraine turned up with some data on a set of Bulgarian made 3.5" disks that had been written on some non western machine.

My first thought was 'nah, it'll have some weird sector map', but it didn't. The disks just read in a standard pc.

And now we're running out of floppies. 

The last production line closed years ago and those people who squirrelled away stock are running out. For the moment you can buy sealed packs of unused disks at an outrageous price, and there are people on eBay who will sell you recycled disks at a slightly less outrageous prices, but it's clear that we are running out of floppy disks.

And while there technical solutions to work round the problem, such as the GoTek floppy emulator, they are of course not certified for use with some devices, and when your device is a Boeing 747, you probably do want to make sure that your emulator has been through the same testing regime as the original floppy drive ...


* actually it wasn't - there was 2.88Mb double density format which never achieved popularity, and which was, from memory, principally found in some 3Com routers and DEC alpha systems ...

Tuesday, 28 February 2023

Lenovo Smart clock gotcha

 I'm a fan of Lenovo Smart clocks.

Minimalist, functional, simple, intuitive, they are essentially a small linux computer which incorporates Google Assistant and emulates all the features of a 1980's clock radio, telling you the time, allowing you to set an alarm, listen to the radio and play music.

It also helps that we got our smart clocks for the cost of shipping. (I had a pile of Telstra loyalty points, and Telstra were selling off their stock of smart clocks.)

Recently I noticed that all of them, and the Google Hub in the study that I use to play podcasts on had taken to telling me it was sunny and 26C. (Actually the Google Hub decided to simply say it was --C).

I wasn't worried, and decided it was most likely a glitch with the weather data feed and that it would fix itself.

It didn't.

Now I control all my Google Assistant based appliances via Google Home running on my Huawei mediapad, which I bought through Amazon AU's marketplace as as a grey market device from Amazon UK.

At the time it was the only recent Android device I owned, so using it to administer the various smart devices seemed sensible.

So I had a look at the Google Home settings.

It said I was in the UK and didn't have a street address. Obviously not correct.

So I changed the location settings to my address in Australia, and immediately the weather settings corrected themselves.

Obviously something must have changed, perhaps as a result of an upgrade to Google Home Assistant, and the application became confused as to location. The fact that I'm running it on a grey market tablet originally destined for the UK possibly didn't help.

So, if you have one of these and the weather details stop working, or are just plain wrong, check the location settings inside of Google Home Assistant ...

Sunday, 19 February 2023

End game at Dow's

 Just before the pandemic struck, I applied for a reader's ticket for the State Library of NSW - I was planning a research trip to search for material relating to some nineteenth century pharmaceutical manufacturers.



Didn't happen - Covid and border closures meant it was a non starter.

But I kept the message at the bottom of my mailbox as a reminder.

Today I deleted the email. 

Covid aside, I reckon that in six or eight weeks I'll be done and have handed over my data before the cold weather makes it unpleasant to work.

I'm not sure what I'll do after but I do know that I don't want to do another big multi-year project ...

Sunday, 5 February 2023

Buffer does mastodon !

 Just a quick note to day that Buffer (https://buffer.com/), which I've used for years to schedule posts on twitter, now works with Mastodon to schedule posts and share web pages.

I've tested it (twice) and it does what it says on the tin ...

Saturday, 4 February 2023

Some more thoughts on Chat-GPT

I've played some more with it, and it's certainly very impressive.

It's remarkably good at producing formulaic text and you can, if you're prepared to put in some effort, get it to produce something clever and amusing.

But that's not the point.

It's simply (a bit of a misnomer here) a very clever text generation engine with a remarkably rich dataset from which to draw it's information. As with all such software, it gives the appearance of intelligence due to the complexity and richness of its responses.

If, for example, I made my living writing real estate advertisements, or some of the formulaic advertorial that goes on real estate sites, I'd be worried - tools like Chat-GPT are admirably suited to that sort of task, but none of its output is so good that it doesn't need to be reviewed by a human.

And that brings me to my second point. 

Chat-GPT's rise to prominence has been accompanied by a wave of  education departments and universities banning its use in essays.

Too late, the horse has bolted.

It's out in the wild, and so it, or some competitor, will inevitably get used. 

So it's time to embrace it.

For example, on my little Chat-GPT exercise on wet plate photography, give the class the example and ask them to critique and expand on it.

I would expect that students would immediately run off and raid wikipedia and camerapedia for information and use that as the basis for an answer. But we're after more than the facts here, we need some critical analysis as well.

Doing so successfully will show that they understand the topic, and it doesn't matter if they use an AI tool to generate the final text - it's the requirement to critique the answer that's important.