Showing posts with label Technology. Show all posts
Showing posts with label Technology. Show all posts

Sunday, March 11, 2018

Protecting copyright with blockchain?

I've been reading articles discussing how blockchain can be used to "protect" the interests of copyright and patent holders.  While I agree this technology would be helpful, we need to recognise that this is a philosophy of "protection" that is the opposite to technological measures such as encrypted media.

Blockchain provides a decentralised database technology, ensuring that records that have been added can't be faked, removed, etc without detection. While blockchain provides a level of authenticity and immutability.of the data not seen before, we are still talking about an enhanced database technology.

I've discussed the flaw in copyright law a few times, which is the outdated interpretation of Berne Article 5 used to claim that there can never be formalities with copyright such as registration.

Blockchain would be a great technology to use, along with modernisation of copyright law, to solve problems ranging from the orphaned works problem to the "not available for sale" problem which I believe is the root cause of a majority of copyright infringement.

Without the modernisation of copyright law, these technologies won't be all that helpful.  The technology would only provide a small benefit for copyright holders who are already visible, while the major problems in copyright law are with works where the copyright holders and licensing options have been kept hidden.

Thursday, August 11, 2016

The Canadiana preservation network and access platform

Canadiana is sending me to Access 2016.  While I expect to learn quite a bit and meet new people, I wanted to make an additional introduction online ahead of time.

I have participated in many changes to the technical platform at Canadiana since I started in January 2011.  As always, what I write is my own thoughts and isn't an official statement from my employer.

Software platform

In early 2011 Canadiana was in transition from an access platform simply called "ECO" for Early Canadiana Online.  This was a mod-perl 1 application that was tightly tied to a legacy version of Apache, was written by an outside consultant, and was in need of upgrading.  I did my best to keep that running as long as possible while keeping our machines secure by having the mod-perl1 application running within a chroot() environment within a newer version of Debian which no longer supported mod-perl1.  This is something we would use linux containers for today, but that wasn't ready in 2011.


The new platform called CAP (Canadiana Access Platform) was also written in Perl, but based on the Catalyst Perl MVC framework, and written by Canadiana staff.  The new platform allowed for multiple "portals" which had their own access theme and content collections (subsets of the full set of AIPs in our repository). CAP used MySQL for most indexes (users and institutions, logging, content metadata) and Solr for search.

Early Canadiana Online remains an important collection, but now shares a platform with other collections.

Between 2012 and 2015 there were major upgrades to the back-end software to allow us to become a CRL certified Trustworthy Digital Repository (TDR).  My part was to upgrade and implement new processes for the automated validation and replication of our file repository.  When I started we were using rsync to copy a comparatively small (A few hundred GB) cmr/ (Canadiana Metadata Repository) directory to a few machines.  The size grew to the point where rsync took a large portion of the day just to realize that no AIPs had been updated. We also ran into a bug in rsync where it can't handle the number of files we were trying to synchronize.

A MySQL based TDR metadata system had been started by a colleague, which I moved to CouchDB to allow the data about our AIPs to be reliably replicated across data centers.  It was now the database that indicted that an AIP needed to be replicated to a specific repository, with rsync only used to reliably transfer all the contents of a specifically identified AIP.  We created processes where a subset of the AIPs would have full md5 checks of their contents, so over time we had validation of every AIP copy that had been replicated to each individual server.

Around this time we also moved from LVM over hardware raid to using ZFS for the TDR file repository, supporting multiple ZFS disk pools individually managed in the TDR file metadata database to optimize storage usage and speed of validation (reduce SAS cabling being a bottleneck).

The new repository system allows us flexibility in how we manage storage, as disk pools can be of any size. A given repository node can be instructed to only store a subset of AIPs as disk space allows, with other nodes having much more storage.  We are not locked in like other systems which only allow storage in a repository up to the size of the smallest node.


In late July 2016 we deployed a major upgrade to CAP to make use of a more Service Oriented Architecture (SOA) for metadata processing, something we call our "metadata bus".  While working closely with the lead developer and our metadata architect, I was the primary on design and implementation of the metadata bus, allowing our lead developer to focus on the upgrades to CAP required to integrate.  While I can write back-end software, I am poor when it comes to front-end and user interface software.

The core of the metadata bus is a series of CouchDB databases, building on the one we created for TDR replication and validation, which we run  microservices against to transform and combine data.  This allows individual microservices to be more easily designed, implemented  and tested than the more monolithic and manual processing we did previously.  This also allows us to easily precompute data which our access platform requires, increasing performance for access.  We continue to use Solr for search, but are now able to harness features we weren't able to with the older revision of the platform.

With this major upgrade fully deployed, I expect smaller incremental upgrades to be deployed more often.

Hardware


In 2011 all our servers were in Ottawa, with some on Wellington where we work and some at a commercial hosting service on Baseline Road.

During the transition from the "ECO" platform we had 2 servers (primary and backup) for ECO and 2 servers (primary and backup) for CAP.  While at the same co-location company, they were in two different cabinets (I only showed one in image, in Green, as this was the publicly accessible server).

On Wellington we had a workflow server (where we processed images from scanning to ingesting into our file repository), a backup server (Black), and a CMR/TDR server (Red)


In 2012 we moved the servers from Baseline Road to a commercial service in Montreal. We had 2 copies of our image data in Montreal (green for publicly accessible, red for standby), and 2 copies in Ottawa (red for standby, and as part of the backup server marked in black).

In 2014/2015 we added a publicly accessible TDR image data at partner University of Toronto, and an additional working copy of the TDR on Wellington (1 publicly accessible copy, 3 standby copies, and 1 copy part of the backup server).


In 2015 we added University of Alberta as a partner, first shipping a copy of the TDR image data and later in 2016 adding an application server.

We also added an application server and a custom project server to our partner at University of Toronto.

This year we separated the hardware configuration for an application server (which needed RAM, CPU and fast disk for database access and search) and the TDR image data/content server (which needed a large amount of storage, and CPU/RAM for making image derivatives). At the same time we designed and documented a WIP (Work In Progress - marked Orange in Ottawa) server to replace the older workflow concepts, and the concept of a "custom projects" server for special things (Marked blue. In Ottawa that houses development and operations tools such as Redmine, Subversion, and Icinga. In Toronto that houses the CHIN Artefacts Canada linked open data project, the legacy Canadiana Discovery Portal which is no longer part of our platform, the one-off Aboriginal Veterans mini-site, and the Drupal-based corporate site we plan to integrate into our platform)



Most recently we decided to move away from commercial hosting providers, and to rely entirely on partners. We picked up the copy of the TDR file repository that remained in the Montreal commercial provider and it waits in Ottawa as we plan for our next partner joining the preservation network.

If you are at a member institution that would like to join the preservation network, please get in touch. While we are interested to hear from any potential partner, ideal is if we can have more provinces represented in the network.  BANQ?  UVic? Memorial University of Newfoundland?

I just spent some vacation time in QuĂ©bec, and would love another excuse to visit UniversitĂ© Laval!  :-)

There are some conversations underway about a new partner, and I am excited to see how things move forward.




Sunday, March 13, 2016

Windows 10 the last desktop version of Windows? The future is unevenly distributed...

I was pointed to a Linux-centric article that included the following section which surprised the person who pointed it out:
Windows 10 will be the last desktop version of the operating system that once gave Microsoft dominance in the PC software market. After that, Windows will be offered on a subscription basis and run from the cloud, but this will not be a Microsoft-exclusive cloud. Internally, Windows will be virtualized within software containers running on Ubuntu.
I'd like to parse this quote a bit, and offer some of my own interpretation.

Last Desktop version

The inevitable disappearance of the desktop operating system has been discussed for decades.  It is really a poor fit for the modern era.  Unlike their more thin mobile counterparts, desktop operating systems really only work well if you have a systems administrator on-hand to handle issues ranging from malware to multi-application compatibility.  System administrators, on the other hand, really want to centrally manage these services so they don't have to spend large parts of their budgets going to each individual desktop to maintain them.  While there is software that attempts to help with this, none of those options can ever compare to running those applications in a server room (local to the office, or in some other server room in "the cloud"). In the server room it is also easier to manage hardware resources, virtualize applications into their own containers to avoid multi-application compatibility issues, and manage software testing and upgrades in a way that is transparent (and thus not disruptive) to users.

Far worse than the problems within medium and large businesses is people running desktop operating systems in small offices or homes that don't have a system administrator.  This why such a high percentage of desktop operating systems are infected by one thing or another -- or just generally not working as well as the hardware and software could work.  This has a high cost to society as a whole, given the harm from spam and malware distribution from this army of infected desktop operating is only surpassed by the fact that these remotely controlled clusters can be bought to be utilized for anything including cyber warfare/terrorism.

I've been looking forward to the death of the desktop operating system for decades, and that is both as a person who works in offices where people expect the IT staff to inappropriately spend a chunk of their budget on desktop support, or as someone constantly asked by less technical family members or friends for help.

It shouldn't surprise anyone that I purchased my wife, mother and my father-in-law each a Chromebook, and as quickly as I could upload all their old documents to a Google drive or put on a USB drive -- and gleefully tossed their old desktops in the trash.

I hope I also will see the eradication of desktop computers in my workplace as well, but that isn't something I have much influence on (even as "Lead Systems Engineer").

What I see in the future isn't less computers, but a recognition that we should be using the right computer and right operating system to fit the job.  The historical "one size fits all" approach that we saw in the desktop era always meant that the operating system used did the job at hand poorly compared to alternatives.

While it is possible that there may be a kernel that will dominate because it receives the most contributions and the most vetting (IE: The Linux kernel), I would consider it yet another market failure if the software stack on top of that remained as similar as we saw in the desktop era.

In my home I run CentOS and Ubuntu on the server, Ubuntu on my development workstation, and we have a variety of mobile devices running Android and ChromeOS.  We have entertainment devices running a variety of OS's (Android on Chromecast, Linux kernel+Boxee software on Boxee box, and a Samsung Smart TV).  While they may all have a linux kernel under the hood, the rest of the operating system built on top is not the same.   I would consider it a backward movement on the part of Google if they merged ChromeOS and Android into the same OS as these two classes of devices serve different purposes and the operating system should be more focused on each purpose.  And I have no interest in running Ubuntu or CentOS on my tablet or phone.

When Google announced ChromeOS they had Citrix there, with the suggestion being in those early days that desktop apps should be virtualized into the server infrastructure, with mobile/portable/disposable devices providing the user interface.

What this will mean is that applications previously run on desktops like office suites and image editing (Photoshop) will be run on servers (in office or in the "cloud") where the computing and system administrators are, and the mobile OS is the user interface only. The desktop application divisions of Adobe and Microsoft have already been moving this direction. The free trials from Google apps may last longer (including still being available free for Gmail users), but they are by no means the only alternative available.

This is also an obvious and long discussed solution to much of the software copyright infringement problem. If you don't distribute software to end users to run in their computers then you don't need to worry about them infringing copyright.   This not only suggests proprietary vendors moving more to the cloud, but that the devices that end users have in their hands will eventually be FLOSS-only.

Subscription?

This is also an inevitable modernization of how proprietary software development will be paid for.  It has never made sense to think of software as a product, as it is more of an ongoing conversation.  While you can buy snapshots of the conversation with a fixed fee, that isn't a useful thing to do when you need to at least keep up with the security patches part of the conversation even if you don't care about new features.

In the early days of computing the hardware advanced quickly as well, and thus people were buying a new computer every few years and thus was paying for new software as well.  Now that computers have reached beyond what the average user needs on their desk/lap there isn't a constant hardware upgrade stream to pay for the massive amount of work that goes into upgrading the software.  In fact, people are wanting to simplify the hardware that they carry with them and want to go mobile where the computing power (as well as battery power consumption) is decreasing rather than increasing.

Subscriptions are the obvious way to go, and this will be of great benefit to both vendors and consumers.

And if you don't want to pay a subscription fee, there will always be legally free FLOSS alternatives. Given software development and system administration time is also expensive this will have to be financed somehow (by someone) even if you never have to pay a software licensing fee again. 

Not be a Microsoft-exclusive cloud

This is also inevitable, and we shouldn't be making a big political deal out of it.  As Microsoft moves away from trying to squeeze percentages out of hardware purchases to being a services company their focus will transition to choosing the right tool for the job.  This will also be a transition away from some of their odd historical political rhetoric in opposition to FLOSS and Linux. Sometimes the right tool will be software from companies and/or open source communities they thought of as competitors in their previous market.

Microsoft's Azure Cloud Switch (ACS) is but one example of this. This isn't a server, desktop, or mobile operating system, but a specialized operating system for network switches built on the Linux kernel.  Using the Linux kernel just makes sense as they can leverage existing community software work, as well as contribute their own code to a community that will then help massively deploy ACS compatible devices.   It is a win-win for everyone involved.

Virtualized within software containers running on Ubuntu

This is the only part of the quote that I'm not convinced was articulated clearly.  Why bother with Ubuntu?   Ubuntu offers a good application server environment, and it works great for workstations, but why bother with the overhead of Ubuntu for a virtualization environment?   This may not be what is being presented in the article.  There may be some value in using the Debian packaging system and build environments, and then spin a virtualization focused distribution.  It might even make sense to build this as a fork of a tiny subset of packages from Ubuntu.  I just don't see using Ubuntu itself as being likely for a company that has the resources to do this right on their own.