2008-07-17

Hilarious - Baby scared of garden hose!

http://www.youtube.com/?v=4JkYEzCJKy8

Hilarious - Baby scared of garden hose!

My poor nephew being picked on by his parents.

2008-05-20

Beta: Rule of Twos

Via Stylegala comes a post by Lois Knight on Typography Essentials. I have been toying with a rule-of-thumb for simplifying web-typography. It's not what I'd call fully tested yet, but it has been working well enough with the students that I've had try it. Here it is:


The rule of twos: No more than two type faces, two colours, two point sizes, two line-spacings and two weights.


Of course you should break this rule, but when you don't have the time, energy or inspiration such rules of thumb can be useful. Oh, it's still in beta so send me your feedback.

2008-05-05

Links for Group Presentation

Faceted Browser examples
Tabulator example Links for my research group presentation. Some examples of web methods for browsing the semantic web.

2008-03-17

Tabulator Firefox extension

I've just started using the Tabulator Firefox extension for browsing the SemWeb. I found that it needed a few tweaks to make it play super-duper nicely.

  • If you have the Piggybank (and Solvent) Firefox extensions then disable these.

Use about:config in the address bar of firefox and change these keys:

  • signed.applets.codebase_principal_support = true (search using the word signed or codebase)
  • network.http.accept.default : add application/rdf+xml to the head of the comma-separated list. (search using the word accept)

Now DBPedia "resource/" links will correctly resolve to using Tabulator, while "page/" links will use DBPedia's own browser. Try it out below:

2008-02-21

Ontological Expressiveness and FOAF

This post on Danbri's Blog highlights an interesting issue in ontological design. Just how specific and expressive do we make an ontology?

TBL has espoused designs that have the "Least Power" which makes them easier to design, implement and use. The sucess of HTTP is a great case in point. As a parallel, in graphic design we have a famous quote about beauty being taking everything out until you can't take anything else out. It's a pretty good philosophy, but like least approached it requires a knowledge of exactly how much is enough and how much is too much. That requires a deep knowledge of how the ontology will be used.

In FOAF, the a single relationship type foaf:knows exists. And most mapping tools I've seen assume a bi-directionality in foaf links - that is if one person lists another as a friend then a back-link is assumed whether or not it actually exists.

Do we want more specificy in foaf relationship types? Perhaps a fuller suite of uni-directional relationship types? In these early stages I'm not sure it's so important to get so specific. Once we get more of foaf:knows relationships then some way to classify these would become more important. For example, Facebook allows an optional level of extra specificity in describing relationship types - that is often necessary given the numbers of friends some people collect in Facebook social networks.

There is some work in this area already: A vocabulary for describing relationships between people

From another perspective, if an enumerated type over a typical dataset forms clusters approaching a single member - and that enumerated type is not meant to be a candidate key and the enumerated label is not naturally a singleton then the enumerated type is probably too specific. (an example of a natural singleton enumeration would be motherof or fatherof.)

And "typical" is the key here. Ultimately it becomes the specialist vs generalist tension that only really gets resolved via de-facto usage.

2008-01-15

SemWeb links of the Week

My Top 3 SemWeb links for the week:
  • DBpedia - community based effort to extract structured data from Wikipedia.
  • Freebase - Community organised ontologies and ontologies with an open api for apps to access
  • Revyu - Review and rate anything

2007-09-29

Generating All Design

I recently had a conversation with my colleague Simon Laing. He suggested that approximately one million by one million pixels is at the limit of human perception, that any more pixels beyond this is not useful. Given that human colour perception is around 16.8 million colours, this means that there exists potentially: 1mil * 1mil * 16.8mil = 1.68 * 10 ^ 19 or 16.8 quintillion visual possibilities.

It's true this would take more than a million monkies; Computers could generate every visual possibility that could ever exist as a snapshot. The computer's owner could try to claim intellectual property rights over every visual possibility that does not already have rights claimed over it. This is of course ludicrous. No computer exists (yet) that is up to the task. And such a brute force approach fails to have any appropriateness context that would make it truly a creative work. Still, it would be an interesting statement to make.

Even if we succeeded in generating all visual snapshots, it would be a useless effort. Such a large number is not infinite, but it is innumerable. Innumerable is a practical infinity. Quite simply there would be little use in having this giant catalogue available since nobody would be able to usefully make sense of it without some form of classification and indexing.

Since generating each alternative is relatively trivial, the real magic is actually in the classification and indexing that allows the space of all possibility to be understood. We view that space using various mental models, world-views and abstractions that zone the space into regions. We layer these viewpoint lenses as a means to make sense of the chaos below. Human language contains ways to describe a visual scene without having to name each 1.68E+19 possibility individually.

Pixels themselves are only an abstraction for representing an underlying objective reality. However, the pixel level is too low for much practical use. For example we speak of web design using higher-level terms such as Banner, Navigation bar, content area, footer, sidebar and column. Like most modes of expression it's about selecting the level of abstraction from objective reality that is most useful to the situation.

This applies to generative design (and generative art). The program code contains the rules that represent the worldview and level of abstraction used to express the design. In generic algorithm parlance the program code contains the rules the transform the genotype (parameters) into the phenotype (output).

And this is only discussing snapshot (human perception is around 25 snapshots per second). What is we added the dimension of time? What about the possibilities of interaction? Photographers also understand the meaning (and thus it's communicative value) is also about the context that the snapshot is viewed in. This is where "beauty is in the eye of the beholder" comes in. Viewers, standing at an exact point in Heroclitus' river form meaning for themselves. Just having the snapshot is not enough without knowing its meaning. This is where the appropriateness measures of creativity come into the equation.

2007-08-31

Monetizing the Semantic Web

Traditional products and service sellers stand to gain customers because the SemWeb will enable location and comparison. But, how do content creators make money selling RDF triplets? Injecting advertising into rich media might make money on the SemWeb. RDF triplets might be filtered at user's computer - and not merely to remove advertising. However, today's computers lack the smarts to remove advertising from rich media; images, sounds, animations, videos, flash interactives. Google has already begun experiments with advertising support for YouTube videos. Micropayments have been proposed since the days of Xanadu and periodically they become fashionable for a time. Clay Shirky's The Case Against Micropayments discusses why cross-site micropayments have never, and will probably never, succeed. Within single sites micropayments do work. Many sites have created internal micropayment currencies. Amazon's S3 storage service uses a micropayment system that bundles many transactions - charged a cent at a time. Bundling is the right idea. Content creators can offer time-limited subscriptions to specific content. A content reseller (the "Info-Vendor) will then aggregate smaller subscriptions into package offerings. This is similar to how content sales already work for online academic journals, TV shows and music. An Info-Vendor service will then become part of the household utility bill - most likely bundled with broadband and cable TV. An Info-Vendor feed may even become a free good subsidized by taxation. This situation will simplify the finding problem of content selection because only a few sources need be queried. The latency aspect of the SemWeb data topography will be reduced resulting in faster and more reliable service. Content credibility can be judged by the Info-Vendors credibility. Ontological translation and basic inference could be run over the Info-Vendor's data store saving local processing time. There is potential for the AI synthetic creation of new content to outstrip human capacity to make use of the information. Info-Vendors may enforce favorite ontologies (and thus implicitly endorse worldview embedded in the ontology). Info-Vendors will also have a controlling stake in the content that users are exposed to. Alternative viewpoints might just not be represented in any of the commercial Info-Vendor's information stores. Contract and intellectual property laws can prevent users on-selling content (automated via proxies), there is difficulty proving ownership of a single RDF triplet. It is just not economical to include DRM at triplet level granularity. Also any DRM systems are at best voluntary. Info-Vendors will gradually lose control over their RDF triplets as the society starts to copy each triplet over and over again; RDF triplets will, in effect, data-leak into the public domain. Therefore, the value of an RDF triplet is in its scarcity. The most successful Info-Vendors will both make available new RDF triplets, including some created using AI synthesis. Like providers of other services, Info-Vendors will ultimately form a varied marketplace. Access to good Info-Vendor service could become the next digital divide. Summary: The ways to make money on the SemWeb are by injecting advertising into rich media and riding the rise of the Info-Vendor.

2007-08-28

Three Content Selection Problems

The three problems of content selection are: finding, filtering and bounding.
Finding has plenty of people working on it. There is Swoogle and (with Web3.0 annotation) even your regular WWW search engine could help.
Once you're found an initial starting point then filtering and bounding are needed. A SemWeb object could have many more triplets describing it then are needed, especially if indirect links are resolved and information is brought in from multiple ontologies. Filtering is the intra-object act of deciding which triplets are important. Bounding is the decisions in how far to traverse the graph of links between SemWeb objects. Bounding is the inter-object counterpart to filtering.

2007-08-24

SemWeb links of the Week

My Top SemWeb 3 links for the week:

Writing Retreat

I had a hugely productive week working on my PhD Research Proposal. Taking a week away from the distractions of work really paid off. Just so my supervisor's know: I've started working all the various documents and scraps into the proper format. It's a mess still, but a big start. My PhD topic is (for now) "Combining Presentation and Content decisions in an Adaptive SemWeb Browser". I have a number of issues that require thinking through:
  • Do I produce something that replaces the current web or works with it?
  • How much do I really care about Web 3.0
  • Should the SemWeb browser be website or desktop hosted?