You know you think about usability too much when...

Last night I dreamed I was hurriedly trying to type something into a cellphone. It had a clever keyboard layout, designed to minimize the number of presses for common letters at the expense of rare ones, so ETAONRISH were just one press each. Fortunately I never had to type Q.

So far, so good. There was just one little problem: it didn't match the keys. They were labeled with the standard layout, where the common letter S takes four presses. The phone worked around this problem by showing the captions on an onscreen keyboard, for the convenience of beginners like me.

Unfortunately the ones it showed belonged to a different clever layout - one that tried to reduce keystrokes by spreading them more uniformly across all twelve keys. (This doesn't make a lot of sense, but I was asleep, okay?) So I was reduced to guessing, while both keyboards helpfully misled me, and the efficient layout was for naught.

Remember those easy, carefree flying dreams? I don't have those anymore. Instead I dream about difficulties controlling my flight. This is what happens when you think about usability problems too much. You start to dream bad user interfaces.

Clojure

While I'm talking about new Lisps, I should mention Clojure, a new Lisp with close JVM integration and fancy concurrency support. I've been meaning to write something in it and post about the experience, but it may be a while before I get around to that. So here's the quick summary:

Clojure has the usual features of modern Lisps: lisp-1-ness, case-sensitivity (because it's hard to do case-insensitivity right), distinguishing () from nil from |nil|, pure symbols, shorter names, modules. It's quite clean, although Lispers should watch out for the renamings (familiar names like cons and do aren't what you expect). There is destructuring in binding constructs, and lambda is really case-lambda. Its Lispy core looks quite good (although as I said, I haven't used it much yet), and it closely resembles the orthodox modern Lisp.

Except for one thing: most data is immutable. Clojure aims at the most popular open problem in languages today: safe concurrency. Like Erlang, it tries to provide state only in forms that are easier to use safely in concurrent programs. Clojure has three: mutable special variables (since their bindings are per-thread), software transactional memory, and reactive message-passing. (See those pages for explanations and examples. There's also a wiki with more examples.) What it doesn't have is the ordinary mutation we take for granted in other Lisps. There's no setf, no mutable arrays or hashtables, no push.

Fortunately there's a lot of careful support for immutable collections, including syntax for [vectors] and {maps}, and generic iteration (even on Java collections!). It includes some useful functions that are often forgotten, such as mapcat and range. Clojure may not have state, but it tries to do the alternatives well enough that you don't feel the lack.

The ultimate alternative to language limitations is an FFI. Clojure is implemented in Java and runs on the JVM, so its FFI takes the form of extensive Java integration: nil is Java null, some interfaces are supported (in both directions!), and it's very easy to call Java code. There are some downsides to staying so close to Java: there's no tail-call optimization, and since threads are Java threads, their number is limited (potentially annoying, in a concurrent language). The upside is that it inherits all sorts of functionality from Java - the VM, massive libraries, even mutable data if you need it. This goes a long way toward explaining the completeness and high quality of the implementation.

There's something I'm wondering about, though. How do you pronounce Clojure? Same as "closure"? Or with /ʤ/ as in "Java"?

Orthodox Otter

Here's another new Lisp: Perry Metzger's Otter is a sketch of a case-sensitive Lisp-1 with short names and a little extra syntax and a module system and regular expressions and a cute animal mascot. Hygenic macros and a fancy compiler are planned but have been put off. It's supposed to be radical and heretical.

Radical? Heretical? Quite the opposite! There are a lot of new Lisps nowadays, and most of them include most of Otter's features (except maybe the mascot). There's hardly anything unusual here, except maybe the distaste for improper lists (a mistake which I also made for a while) and the choice of ! for setf (which seems obvious in retrospect - I will steal it now kthx). The rest of Otter is a reflection of what most Lispers consider the right way forward. It's not heretical, it's orthodox!

It's hard to detect these changes from the inside, but it seems to me this is a new development. Ten years ago, there was not so much agreement on what new Lisps should look like, nor on their desirability. The old orthodoxy was to cling to the existing dialects, as the language was perceived to be under threat. Remember when everyone was defending Common Lisp, except the Schemers, who claimed their language wasn't even a Lisp? But nowadays most Lispers don't think of the language as threatened. Instead we look forward to its future growth, and we aren't afraid to let a hundred flowers bloom.

But Otter probably won't be one of them. That presentation, from last August, seems to be the only available information on it. There's still nothing on the website but a picture of some sea otters. Sketching new lisps is fun (well, I think so, and evidently Perry does too) and easy, especially when you have an orthodoxy to guide you. Implementing them, and solving all the problems that surface in the process, is a lot of work, and prone to being put off forever. I suspect that's what happened to Otter.

But where one person can sketch a Lisp, others can build one. And we should. Never in the history of Lisp (or any programming language, IMO) have we had such an appealing orthodoxy. It is an irony not surprising in the Lisp community that that orthodoxy is now a language that does not actually exist; you cannot use it. But it is the sort of language that makes you want to. The popularity of new Lisps, and the extent of agreement on their features, suggest that it will have many incarnations in the years to come.

And now I'm better go work on mine...

For want of a rant

For my day job, I hack C++. A great many rants could begin this way, and after a week of long days debugging crashes in deployed software, I'm tempted to write one. Why do so many programmers insist on using a low-level language for everything, whether it's needed or not?

But you've heard that rant before, and maybe even written it yourself; you know how it goes. So I won't bother. Besides, I'm not sure I agree with it. It's easy to blame any problems with C++ on the language's low level, but is that really where they come from?

One of the crashes I encountered this week was a buffer overflow, due to careless use of strcat. A classic case of C++'s unsafe, low-level data structures, right? Except that C++ does have real strings, and there was no reason not to use them. We just didn't use the tools the language gave us.

Part of the application I deal with at work is in Java, and crashes in that part are never hard to debug, because they always come with handy stack traces. The problem with C++ isn't that it crashes - that could hardly be prevented. It's how little information it reports when it crashes. But this is not C++'s fault. The information (except for symbols) needed for a stack trace is already there in the crashing image; it's just not normally reported. One must go to the trouble of using a debugger to extract it. This is a place where a small change in an operating system could make a big difference in a language's usability. On a system which generated stack traces by default when native code crashed, C++ would not be nearly as hard to debug.

Of course C++ is unsafe, and therefore sometimes clobbers the necessary information (as it did in that buffer overflow). And its limited expressiveness often means more code and more bugs than in a higher-level language. But not all of its troubles are a result of low language level. And some of them aren't even its fault.

For my day job, I hack C++. And it's really frustrating. It doesn't even give me a good excuse for a rant!

Taboo

Most arguments are not a result of failure to communicate, and most communication problems are not a result of vaguely or multiply defined terms. But for those that are, Eliezer Yudkowsky suggests a solution: declare the poorly defined word taboo, and try to communicate without it. This is obvious, of course (at least in retrospect), but Eliezer points out what is not so obvious: avoiding a word can be useful even in the absence of an argument, to expose confusion that is masked by familiar terms.

Maybe it's just an illusion of familiarity, but it seems to me that the subject of programming languages has more than its share of such confusion. We talk about formal, unambiguous languages, but the language we do it in is rather sloppy. Many important words are defined so vaguely, or have so many disparate meanings, that they serve less to convey information than to encourage the listener to imagine whatever meaning they please. This is not a huge problem when talking to someone who knows what you mean, but between people with less shared context it's a fount of unenlightening arguments. So I think Eliezer's taboo trick could be useful.

Take object-oriented, for instance. It has a wide range of meanings, so regardless of which one you mean, you can be reasonably confident that your audience will misunderstand. And it's emotionally charged, so those misunderstandings will become arguments. Better to avoid it. You might fall back on words like encapsulation and polymorphism - but both of those are also ambiguous, and in ways that conceal controversies! Clarity is harder than it sounds.

Of course it would be unbearably tedious to say everything with rigorous precision. But much of our vocabulary is pointlessly vague, and I think it creates a lot of confusion with no gain of convenience. So take this as a challenge. Can you say what you have to say about programming languages without using those common words that have no common meaning? How much can you say without mentioning syntax, or functional, or type?

RSS is a dangerous drug

I didn't understand the point of RSS until I started using a feed reader about six months ago. Then I saw what I'd been missing. It automates something I did not realize I spent effort on: keeping track of what I haven't read yet. Doing that by hand - visiting sites periodically, having lots of tabs open, trying to remember what I haven't looked at recently - is time-consuming and unreliable. So it's a perfect candidate for automation. Handing it over to a machine means I don't forget to check a site, or lose track of something when I get distracted and don't finish reading. It means I can be confident I won't miss anything. It means I can read even more blogs!

Sometimes this is not a good thing. I often feel compelled to read every post of every feed I subscribe to, and I have to remind myself otherwise. And I'm always tempted to subscribe to anything that looks remotely interesting. But if I subscribe to more feeds than I can handle, or don't read them for a while, the unread posts pile up. I haven't checked my feeds for the last week, and I now have 528 items to read. For a compulsive reader like me, an ever-growing list of interesting things to read is dangerously addictive. Fun, yes. But tyrannical.

OK, it's down to 348 now. But this feels like work!

Multiple kinds of type in one language

It's not just different languages that distinguish types in different ways. It's pretty common for one implementation of one language to have several type mechanisms serving different purposes. For instance, a typical optimized Lisp will have:

  1. Multiple representations of variables (integer, float, and tagged pointer), for efficient unboxed arithmetic
  2. Tags distinguishing pointers from fixnums (and sometimes characters, nil, and immediate floats)
  3. Dynamically typed objects
  4. A sublanguage for describing sets of objects, especially inferred or declared static types.

The first two are not called "type" in Lisp, but they would be in a lower-level language. The last two are both called "type" even though they are quite different things. Perhaps surprisingly, this doesn't cause much confusion in practice.

An ML implementation might have:

  1. The same multiple representations of variables
  2. The same tagged pointers (they make garbage collection simpler, even if they're not actually used)
  3. Dynamically tagged objects, to allow sum types
  4. Those famous static types, which are the only one of these mechanisms that's normally called "type" in ML.

The first three mechanisms are basically the same as in Lisp, despite the languages' seemingly opposite approaches to type. Where they differ is in how they describe types - and of course in what static types they allow. I find it particularly amusing that ML sum types use the same mechanism as dynamic type. It's not called that, and it's not used the same way, but ML data is partly dynamically typed.