Saturday, January 10, 2009

Doing Science in the Sink

I just performed an interesting and unintentional scientific experiment with an empty milk carton. After finishing off a gallon of (skim!) milk, I went to rinse it out so that the milk didn't sit and smell. When I first turned the faucet on, it was set to spew out hot water, which I realized was unnecessary, so I changed it to cold water after a few seconds. I put the top on, and shook.

When I went to open the carton, I noticed that the sides were puffed out and when I took the top off, a bit of air rushed out. Interesting, I thought, why does this happen?

After a few tests, I found that I was reliably able to reproduce this effect by filling the carton with a few seconds worth (around 3ish) of the hottest water the sink could put out, switching to the coldest water it could put out, and then putting the lid on and shaking. Every time, it would puff out the bottle. Also, filling it with just hot water or just cold water did not have an effect.

After spending a few minutes testing, I had isolated the effect, but did not know the cause. Luckily, my dad happened to be around, and suggested that the answer had something to do with the fact that cold water has a higher capacity for dissolved gas than warm water.

It is just a guess, but I think it could be that the cold water brings extra oxygen (or whatever is dissolved in it) into the carton, and when it is mixed with the hot water, its ability to retain the dissolved gas decreases. The warm and cold water stay separate long enough that there is time to put the cap on, and shaking the bottle mixes them up and causes the cold water to release its gas.

If this is correct (and there is an excellent chance that it is not; I haven't taken chem since 10th grade), it would also mean that the cold water gives up more gas by heating up than the hot water absorbs by cooling down. This seems to suggest that the gas absorption to temperature function is non-linear (or my measuring methods of "about 3 seconds" are not very precise).

Try it at home! All you need is an empty milk carton (or any largeish container which you can close quickly), a faucet which can switch between hot and cold water fairly quickly, and your imagination!

Tuesday, January 6, 2009

Interesting Look at the Financial Meltdown

I saw this article sitting on the kitchen table the other day and then set about to find it on the internets.

It is fascinating to see how badly the whole financial system was set up that a guy, Harry Markopolos, could spend 9 years specifically detailing how Madoff's company could not be anything other than a Ponzi Scheme, and be basically ignored by the regulatory bodies whose job it was to investigate exactly those issues.

His argument went something along the lines of:

  • Madoff claims to be handling roughly $50 billion in securities.

  • The markets in which he claims to be involved are smaller than that.

  • If he actually did put $50 billion into the market, we would see those trades somewhere.

  • We don't see them, so he can't have invested the amount of money he claims.

The response seemed to be of the form, "yeah, we'll look into it." Turns out Markopolos was right, and nobody is quite sure where that $50 billion went.

The real meat of the column, however, is how the Madoff scandal (and the financial crash in general) was caused not so much by greed, but by a breathtakingly stupid system. Evie would be proud.

Take a look at the article for a better explanation, but the basic problem was that the people running the SEC were using it as a stepping stone to get a job at the same places they were supposed to be regulating. So if Bob the SEC Director of Enforcement was really hard on, you know, criminals who conned people, when he went to them asking for a job, they would probably roll their eyes at "Bob the Buzzkill." On the other hand, if he turned a blind eye, or, even better, actively helped them steal more money, come hiring day, he would be "Bob, the cool guy who got us all those strippers and booze," a much more qualified candidate.

It seems like the solution will come in the form of a system which does not reward and encourage the exact opposite of what it is intended to achieve. How to do that, especially in the face of an established system which, until recently, was doing quite well for itself is an exercise left to the reader.

On the Baddity of Cities

I told you all! Cities are bad for you!

I have heard some criticism of the theory that the article presents, that the hypothesis is not really supported by the data, or that the fact that simply looking at pictures of nature improved performance on certain tests means that something else could be at work. I think these are valid, and that the topic needs some more research, but it is an interesting and intuitively correct hypothesis.

If I spend more than a few days in a city, I do tend to get physically tired and overwhelmed. Each time I go to New York, I am reminded of this. I find that everywhere I go, I am constantly scanning the surrounding area for threats and information. After a few minutes waiting at a subway stop, I have looked at the face of everyone in the terminal, and if anyone new comes in, I turn, look at them, and go back to waiting. It is subconscious, but it causes a lot of stress, since my environment is constantly changing and I am struggling to keep up. If I focus on ignoring my surroundings, I can avoid looking around, but that takes concentration and I still feel uneasy when I hear a turnstile turn but don't look at who came through.

One of my favorite parts of Andover was being able to take a short walk from my dorm and head to the ~100 acre sanctuary, which was a big forested area with a looping path going through it and two ponds. It was my favorite place to go jogging, and was a great way to relax from the pressure of the place. There was rarely anyone else there, or at least it was big enough that I didn't run into people very often, so I could just concentrate on running. In contrast, I haven't found anywhere at Brown where I enjoy running, since I hate running in the city. I tried it a few times, and it was just unbearable. I have to be looking out for cars and people, there are things and places to avoid and I have to remember my route back. If I'm running, I don't want to have to think about anything, and I can't do that in a city (even Providence).

So this article completely made sense, and I would like to see more studies to better establish the validity of the claim for a broader range of people. For me at least, I have known this for a long time.

On the bright side, it could be a lot worse.

Sunday, January 4, 2009

To Cage an Idea

This is a reply to a New York Times article sent to me Evie entitled Who owns your great idea.

I saw a mention of the iShoe guy earlier today, but didn't read too much about it, so it was nice taking a closer look. My initial reaction is that MIT is being overly stingy with the licensing and ownership, and it is creating a worse situation for research. In the true spirit of Laissez-faire economics, the "research-by-students" market is readjusting itself in that, according to the article, students burned by this policy are moving their operations off-campus.

The problem I see with it is that it becomes a huge duplication and waste of effort. If a university sets up a bunch of labs, machine shops, and facilities for its students to use, but they all go off-campus to be able to keep ownership of their inventions, all that lab space at school is wasted.

I can see the school wanting to make back money on its investments in facilities and teaching, both of which can be hugely expensive. However, it seems like there should be a clear distinction between a "university" and a "corporate R&D department." To me, the purpose of a university is to promote the progress and learning of its students as well as support its faculty to conduct basic research. The R&D department is devoted to creating new products for its company to sell, with a focus on shorter-term commercialization.

These two goals are not incompatible and often intersect, but some level of separation should be kept between them in order to let them do what they do best. Universities can conduct research on a wide range of topics without the pressures of needing to produce a viable product. On the other hand, R&D labs have the task of making something you can actually use without a PhD in mechanical engineering. Universities produce tons of raw ideas, companies mine that raw idea-ore and smelt it into products.

A school saying "we want to try to commercialize anything coming out of our students" is like a group of civic engineers saying "let's forget about that whole bridge-building thing and just do some quantum physics." If you want to do quantum physics, become a quantum physicist, but in the mean time, people need bridges now. Similarly, if you want to make money off of selling products, become a company, but don't interfere with the learning process in order to make a buck.

If a student has money to develop and protect an invention on his own, should he? It depends. Turning the job of commercializing a product over to a university is a better bet if you have no interest in business or if finishing school is a priority.


That's a false dichotomy. The student could put the research in the public domain, have a nice bit of padding on their resume, and let the research benefit the maximum number of people.

Invention-hoarding (or "intellectual property" in general) seems to be good[1] for inventors, bad for inventions. It isn't pretty hard to see that being able to sell licenses for an invention you make gives you more money. On the other hand, Erez Lieberman (the iShoe guy) now has to pay $75,000 up front to MIT just to try to commercialize his idea (and even more should he succeed). Also, assuming he does gain exclusive rights, what if someone else wants to come along and make the same thing? If the idea were unpatented (or whatever the equivalent to public domain is for patents), then anyone could try their hand at making such a shoe, and whoever had the best design or best business plan would succeed. If the license is exclusive, only one person can work on a self-balancing shoe at a time.

On a brief societal note, the challenge would become to figure out an appropriate balance between good for the inventor and good for the invention such that a maximum number of stuff is created. Too much "intellectual property" ownership and individual ideas are never combined due to prohibitive licensing costs; too little, and nobody ever makes enough money off of their inventions for it to be worthwhile to invent them (in theory. It's assumed that this would be the case, so nobody has ever really tried).

Last month, the university determined that while the students own the design, R.P.I. owns the idea for the bottles. The students must license it from them [...]


What? That's just absurd. I have an idea for a machine to solve world hunger. Now anyone who invents one may own the design, but they have to pay me for the "idea." How is this "promoting the Progress of Science and useful Arts" again?

"At the same time, they have real potential, and our goal is to encourage them."


That seems to be the heart of it (right at the end of the article). By maintaining strict control over ownership of ideas and design, the university is discouraging invention with university resources, making it more difficult to conduct research and development at all. It seems like RPI is coming up with a more sensible policy, so I guess some good is coming of it.

1. That part isn't even completely clear-cut. A company like Red Hat wouldn't be viable if it didn't share alike its inventions.

Tuesday, October 7, 2008

On Linux on the Desktop

This is a message I sent in response to an email from my dad. He was commenting on the recent story about the high return rates for Linux Netbooks.

I was reading more about this company that was selling the cheap Linux notebook, and getting lots of returns.


I think this was in reference to Netbooks (not "notebooks," these are smaller and lower powered than a traditional laptop) being sold by MSI. MSI is primarily a hardware company that targets OEMs and people who would assemble their own computers (like me) and so doesn't have much experience making usable operating systems.

Apparently, they (foolishly) decided to use a custom Linux distribution rather than something like Ubuntu. It seems like the problem is more one of MSI giving people a bad installation rather than an inherent problem of Linux.

This particular article was pointing out that people who bought this in the first place were more adventurous and more knowledgeable than most computer users to begin with, but still could not deal with it.


Again, I think that this was largely a problem of a poor install than a fundamental problem of Linux.

I would amend to say "Linux is a great choice for extremely technically sophisticated users who prefer being as far as possible from the mainstream."


That has certainly been the traditional user-base, and is still a significant part of the development community (Richard Stallman refuses to use the *Web*), but there is effort going into changing that. With the increasing popularity of Linux on these Netbooks (would this story even have been possible a few years ago?) as well as cell phones, there is a lot of effort going into usability improvements.

The great thing about open source programs is that it is very hard for useful programs to "just die." If a commercial program loses its corporate overlord, it can fade out and whither away. If a company gets bought up or out-competed, applications can disappear. This has been the Microsoft strategy. The reason they are so scared of Linux and open source is that even if you kill every developer of open source programs, the code is still there, and anyone with the knowledge and inclination can work on it.

Since most open source programs don't have the burden of needing to make money off of their direct sale, they tend not to get worse for the sake of adding features. Look at any version of Norton after around 2003. They needed people to keep buying the program, so they needed to add /something/ to make it different from the previous version. The problem is that it basically already did most of what it needed to do, so they had to add un-features and made it worse than it was, to the point where a computer was better off without Norton than with it.

For open source programs, if they reach maturity, people will maintain them, but not add features for the sake of making more money, since the developers generally don't make money directly off of the sale of the program. This means that programs which are basically done don't try to add useless features.

Also, it seems like investments in technologies and frameworks provide more of a network-effect benefit within the open source world than in the proprietary world. Open source has been playing catch-up for a while, but it is starting to pull ahead, with the web browser space is the most dramatic example. While it took a while for Firefox to reach parity with Internet Explorer, the current version of Firefox is much faster and more featureful than the latest IE, and the development version of both Firefox and WebKit (the engine that powers Safari) have javascript execution engines up to 40x faster than IE.

The level of polish and development that constitutes "acceptable" is not static, but it is not moving as fast as the development of the open source ecosystem. Ubuntu is usable for many people's day-to-day tasks already, and is only getting more usable. As time goes on, it will become acceptably easy for an increasing number of people.

Linux can take over from Windows, but they need to make it easy. And for that they have a ways to go.


Agreed, but there has been huge progress within the last few years. And it shows no signs of slowing.

It is sort of like in my field we generate many different kinds of images. The neurologists complain that the labelling of the images is inconsistent so "they can't tell what they are looking at". We never look at the labels, because it is obvious from a glance what kind of image it is. So to us an elaborate system to produce consistent labels would be so useless as to be a waste of whatever time it took to implement. To the neurologists not having it is a problem. If I were as into Linux as I am into brain images, then I suspect I would find GUI as useless as you do. As it is, I have fewer demands on my computer, but "labor saving" is at the top of the list.


I think this is what has kept the GUI less newbie-friendly than that of Mac or certain aspects of Windows (I would assert that Ubuntu is more user-friendly in many aspects than Windows, but not as familiar to many people). It is easier for a commercial company to hire usability experts and compel interface designers to produce good GUIs than for a group of hobbyists who don't mind a CLI to spontaneously produce a good GUI.

The instructions may not always be clear, but they do not run to "type the following with exactly this syntax, except, of course, changing the part you need to change."


I will have to say that it is a lot easier to tell someone:

Type this:

ps aux | grep zfs

and copy and paste the results to me


than it is to say,

Open task manager, try to find each entry with the string 'zfs' in it, and tell me what appears in each column.


On the other hand, it is a lot easier to click the "Applications" menu and then look through the "Accessories" or "Games" menu rather than memorizing that your chess game is launched with "glchess" or that "Manage Passwords and Encryption Keys" is launched with "seahorse." It all depends on what you're trying to do and how much up-front time you are willing to invest in order to save time later.

Saturday, June 7, 2008

Really Really Simple Syndication

From the everything is miscellaneous and web that wasn't talks, it seems like what we really need is some way for computers to figure out what information we want and give it to us. There is a huge amount of info encoded in blogs, the example of nearly one blog post for each word in a Bush speech about immigration, but so far, there is no good way of extracting much of it. Sure, we can google for "blogs about bush's immigration speech," but even that would likely turn up a bunch of blogs about random things, a bunch of pages about unrelated immigration topics, and so on.

The first step would be to be able to do something like "tell me all of the instances in which bush has changed his position on immigration" which would look at all the blogs which referenced the speech and found things relating to a change in the stated position. This would probably require natural language understanding of queries, but that might not even be necessary, as it could simply relate the words in the query with the words in the blogs.

The next thing needed would be for it to be anonymous (or at least have that possibility) and distributed. Aside from privacy concerns, there could be massive scalability problems, a la Twitter. There is the concept of a distributed hash table, but that only works for exact matches. Wikipedia to the rescue, with this (PDF) bit of magic from Berkley.

Personalized feeds



However, this would only be the beginning. Once you can respond to general natural-language queries like that, you could build up a list of what someone was interested in. RSS feeds are a fascinating concept and allow people to get a feed of news customized for their tastes, but it assumes that everything at a certain address will be interesting to me.

If you go to the RSS page on the New York times, you will see that they provide feeds for "Business," "Arts," "Automobiles," and so on. Going down further, there are feeds for just "Media & Advertising" "World Business" under business and "Design" and "Music" for arts, but nothing under automobiles. This is the kind of problem that David Weinberger was talking about: what if all "Arts" are pretty much the same for me, but I want to differentiate between "Foreign Cars" and "Domestic Cars," or even "Sports Cars" and "Trucks." Maybe I don't care about red cars at all, so I don't ever want to see a story about red cars in my feed.

The problem with RSS is that, even though it is a huge step forward for allowing us to keep track of many different sources of news at once, someone else has to decide how to split up the feeds. Most places give you an option of what you might want; the current system the New York Times is much better than having one giant "this is everything" feed, but it still has a ways to go. This could be done without having to rewrite our RSS clients by allowing the user of a site to set up a custom RSS URL from which to pull updates, but this would place a huge burden on the content providers and would not scale at all.

The solution



So after we have a good way for the user to tell the system what to get, we would want a way for the system to learn what the user likes, possibly using a Netflix-like recommendation system, and automatically pull down stories that the user likes.

If President Bush gives a new speech about policy for the Internets and 1,000 people blog about it, our little system should automatically sift through all of the posts, figure out which parts of the blogs would likely interest you, based on what you have told it you like in the past, and present you with some sort of reasonable compilation of the information, all without any user interaction. This would really unlock the potential of the Internet. Get hacking.

Thursday, May 29, 2008

Missing the point

While making my daily rounds of programming language-related blog posts, I came across a couple items which caught my attention. I have been very interested in parallel and concurrent programming of late, especially how to solve the big issues everyone seems worried about relating to the non-easily parallelizable code.

I saw a couple of posts, after following a couple of branches off of a Slashdot story, which seemed to confuse a few of the issues surrounding parallel programming. This one in particular, confuses a number of different problems.

Starting off, he correctly points out that:

Users don't care about parallel processing anymore than they care about how RAM works or what a hash table is. They care about getting their work done.

Assuming that he is not talking about "people writing programs" as users, he is absolutely correct. As long as something works well and fast enough, nobody cares.

But therein lies the problem: well and fast enough. The "well" part is fairly simple: if you are doing multiprocessing, it still has to work. That's pretty obvious, and while it can be challenging at times, there is no real controversy over that fact.

This leaves the "fast enough" part. The problem here is that since the dawn of time (which according to my computer is January 1, 1970), people have been able to count on future computers getting faster. Moore's law and all. Nowadays, computers get faster by adding more cores, but software is still written assuming that we will get this hardware speedup. The problem is that the hardware guys are tired of giving the software guys gigantic speed boosts without forcing the software guys to change their behavior at all. They are holding a worker's revolution and throwing off the shackles of the oppressors and saying, "you want more speed, write your programs more parallely!"

He and the guy here do mention the problem of playing music while surfing the web and checking email (which are often both in the same program, but whatever). On a single core system, playing a song, unzipping a file, and running a browser would lead to some whole-system slowdown, and this problem was greatly helped by multi-core computers, nobody is arguing that. The problem is that this process-level parallelism only helps you until you have more cores than processes, which really isn't that far off. Think about how many programs you are running now. Aside from OS background services which spend the vast majority of their time asleep, there are probably between 2 and 5.

Let's say that you are playing a movie, browsing a flashy website, talking to someone over VoIP, and compiling your latest project. That's four processes. You will see great improvements in responsiveness and speed all the way up to 4 cores, but after that, you will be sitting idle. If each of those programs is single-threaded, or its main CPU-intensive portion is single-threaded, then adding more cores won't help you at all.

This is the point that the author of the second article misses. There may be 446 threads on his system, but, as he observes, many of them are for drawing the GUI or doing some form of I/O. Drawing a GUI takes relatively little processor power compared with something like encoding a video (unless you are running Vista, har har), and for I/O threads, most of what they are doing looks like this:


while (! is_burnt_out(sun)) {
data = wait_until_ive_gotten_something();
copy(data, buffer);
}


In other words, not really doing anything with the CPU. This means that, although there are a great number of threads, only a few of them (Firefox does all of its rendering and JavaScript execution in a single thread for all open windows and tabs) actually will be using the processor. The "crisis" that everyone is talking about is when those single computation threads start to get overloaded. With so many things moving to the browser, what good is a 128-core CPU if all of my precious web apps all run in Firefox's single thread? I have these 768 cores sitting around, wouldn't it be nice to use more than 2 of them for video encoding?

Just Make The Compilers Do It



One thing that the second articled brings up is automatically-parallelizing compilers. I do think that there is something to be said for these, especially since it has been shown over and over, first with assembly and then with garbage collection, that compilers and runtimes will get smarter than programmers at doing things like emitting assembly or managing memory. I would not rule out the chance that a similar thing will happen with parallelization.

I do think that making parallelizing compilers will not be as "easy" as writing good optimizing compilers or good garbage collectors. The problem is that the compiler would have to have a broad enough view of a serial program to be able to "figure out" how it can be broken up and what parts can be run concurrently to be able to generate very concurrent code. This would go beyond inlining functions or unrolling loops to figuring out when to spawn new threads and how and when to add locks to serial code to extract maximum performance. Far be it from me to say that we will "never" come up with something that smart, but I seriously doubt that we will be able to code exactly as before and have our compiler do all the magic for us.

Save Us, Oh Great One!



So what do we do? The answer is not yet clear. There are clearly a lot of problems with traditional multithreading with locks (a la pthreads and Java Threads), but nobody seems to agree on a clear better way of doing things. I saw a cool video here about a lock-free hash table. The idea of using a finite state machine (FSM) for designing a concurrent program is fascinating, but I could see problems with data structures involving more dependent elements, like a balanced tree. Still, I think the approach gives a good insight on one way to progress.

A similar, but less... existing idea is that of COSA, a system by renowned and famous crackpot Louis Savain. He uses an interesting, but basically unimplemented graphical model for doing concurrent programming and talks about the current problems with concurrent code.

Now, judging from the tone of this article and the titles of some of his other posts (Encouraging Mediocrity at the Multicore Association, Half a Century of Crappy Computing), he seems to be 90% troll, 10% genius inventor. I took a brief look at his description of COSA, and it seems to have some similar properties to the (much more legitimate) lock-free hash table. The idea of modelling complex program behavior in a series of state machines or in a big graph, but a lot would need to be done to make this into a practical system.

As interesting as COSA looks, the author seems to be too much of a Cassandra for it to gain any appeal. I mean "Cassandra" in the "can see the future, but nobody believes him," sense, not the "beauty caused Apollo to grant her the gift of prophecy" part (but who knows, I've never seen the guy).

I see some evolution of parts of each of these ideas as being a potential solution for some of the multiprocessing problems. The guy who made the lock-free hash table was working on the level of individual CASs, which would definitely need to change for a widely used language. Just as "goto" was too low-level and was replaced with "if", "for", and "map", some useful abstraction over CAS will come along and enable concurrent programming at a higher level of abstraction.