<?xml version="1.0" encoding="UTF-8"?>
<?xml-stylesheet type="text/xsl" media="screen" href="/~d/styles/rss2full.xsl"?><?xml-stylesheet type="text/css" media="screen" href="http://feeds.feedburner.com/~d/styles/itemcontent.css"?><rss xmlns:feedburner="http://rssnamespace.org/feedburner/ext/1.0" version="2.0">
   <channel>
      <title>Common Knowledge</title>
      <link>http://scienceblogs.com/commonknowledge/</link>
      <description />
      <language>en</language>
      <copyright>Copyright 2010</copyright>
      <lastBuildDate>Mon, 22 Mar 2010 19:09:52 -0500</lastBuildDate>
      <generator>http://www.sixapart.com/movabletype/?v=4.32-en</generator>
      <docs>http://blogs.law.harvard.edu/tech/rss</docs> 

      
      <atom10:link xmlns:atom10="http://www.w3.org/2005/Atom" rel="self" type="application/rss+xml" href="http://feeds.feedburner.com/scienceblogs/CommonKnowledge" /><feedburner:info uri="scienceblogs/commonknowledge" /><atom10:link xmlns:atom10="http://www.w3.org/2005/Atom" rel="hub" href="http://pubsubhubbub.appspot.com/" /><feedburner:emailServiceId>scienceblogs/CommonKnowledge</feedburner:emailServiceId><feedburner:feedburnerHostname>http://feedburner.google.com</feedburner:feedburnerHostname><item>
         <title>Open Hardware</title>
          <description>&lt;p&gt;Creative Commons was fortunate enough to be involved in a &lt;a href="http://eyebeam.org/projects/opening-hardware"&gt;fascinating workshop&lt;/a&gt; last week in New York on Open Hardware. Video is at the link, photos below.&lt;/p&gt;

&lt;p&gt;&lt;object width="400" height="300"&gt; &lt;param name="flashvars" value="offsite=true&amp;lang=en-us&amp;page_show_url=%2Fgroups%2F1363786%40N20%2Fpool%2Fshow%2F&amp;page_show_back_url=%2Fgroups%2F1363786%40N20%2Fpool%2F&amp;group_id=1363786@N20&amp;jump_to=&amp;start_index="&gt;&lt;/param&gt; &lt;param name="movie" value="http://www.flickr.com/apps/slideshow/show.swf?v=71649"&gt;&lt;/param&gt; &lt;param name="allowFullScreen" value="true"&gt;&lt;/param&gt;&lt;embed type="application/x-shockwave-flash" src="http://www.flickr.com/apps/slideshow/show.swf?v=71649" allowFullScreen="true" flashvars="offsite=true&amp;lang=en-us&amp;page_show_url=%2Fgroups%2F1363786%40N20%2Fpool%2Fshow%2F&amp;page_show_back_url=%2Fgroups%2F1363786%40N20%2Fpool%2F&amp;group_id=1363786@N20&amp;jump_to=&amp;start_index=" width="400" height="300"&gt;&lt;/embed&gt;&lt;/object&gt;&lt;/p&gt;

&lt;p&gt;The background is that I met &lt;a href="http://eyebeam.org/people/ayah-bdeir"&gt;Ayah Bdeir&lt;/a&gt; at the Global Entrepreneurship Week festivities in Beirut, and we started talking about her &lt;a href="http://littlebits.cc/"&gt;LittleBits&lt;/a&gt; project (which is, crudely, like Legos for electrics assembly - even someone as spatially impaired as me could build a microphone or pressure sensor in minutes). &lt;/p&gt;

&lt;p&gt;Ayah introduced me to the whole &lt;a href="http://en.wikipedia.org/wiki/Open-source_hardware"&gt;open hardware&lt;/a&gt; (OH) world and asked a lot of very good, hard to answer questions about how to use CC in the context of OH. It became clear that a lot of the people involved in the movement didn't have a clear grasp of how the various layers of intellectual property might or might not apply.&lt;/p&gt;

&lt;p&gt;Ayah suggested in February that we put together a little workshop - almost a teach-in - around a meeting of &lt;a href="http://www.arduino.cc/"&gt;Arduino&lt;/a&gt; advocates happening in NYC on the 18-19 of March. In a matter of three weeks, we got representatives from a bunch of major players to commit: Arduino (world's largest open hardware platform), &lt;a href="http://www.buglabs.net/"&gt;BugLabs&lt;/a&gt;, &lt;a href="http://www.adafruit.com/"&gt;Adafruit&lt;/a&gt;, &lt;a href="http://www.chumby.com/"&gt;Chumby&lt;/a&gt;, &lt;a href="http://makezine.com/"&gt;Make&lt;/a&gt; magazine, even Chris Anderson. Mako Hill from the Free Software Foundation came and @rejon made it there at the last minute too, wearing his &lt;a href="http://wiki.openmoko.org/wiki/Main_Page"&gt;openmoko&lt;/a&gt; and &lt;a href="http://en.qi-hardware.com/wiki/Main_Page"&gt;qi hardware&lt;/a&gt; hats. &lt;a href="http://eyebeam.org/"&gt;Eyebeam&lt;/a&gt; hosted it for free, and we picked up the snacks and cheese trays.&lt;/p&gt;

&lt;p&gt;I gave a &lt;a href="http://www.slideshare.net/wilbanks/open-hardware-briefing"&gt;very short intro&lt;/a&gt; laying out how the &lt;a href="http://sciencecommons.org/"&gt;science commons project&lt;/a&gt; @ &lt;a href="http://creativecommons.org/"&gt;creative commons&lt;/a&gt; has spent a lot of time looking at IPRs as a layered problem, dealing with it at data levels, materials levels, and patent levels, as well as the fact-idea-expression relationships in science. This was to create some context for why we might have interesting ideas.&lt;/p&gt;

&lt;p&gt;Thinh proceeded to deliver a masterful lecture on IP that went on for hours, though intended to be 30 minutes. It was an interactive, give-and-take, wonderful session to watch, ranging from copyrights to mask works to trade secrets to trademarks and patents. The folks there liked it enough to suspend the break period after five minutes and dive back into IP.&lt;/p&gt;

&lt;p&gt;After that we had a lengthy interactive session driven by the OH folks in which they tried to decide what a declaration of principles might look like, how detailed to get, how to engage in existing efforts to do similar things (like &lt;a href="http://www.ohanda.org/"&gt;OHANDA&lt;/a&gt;), the role of the publishers like Wired and Make to support definitions of open hardware, and how open one had to be in order to be open.&lt;/p&gt;

&lt;p&gt;There was no formal outcome at the close of business, but I expect a declaration or statement of some sort to emerge (akin to the &lt;a href="http://www.soros.org/openaccess"&gt;Budapest Declaration on Open Access&lt;/a&gt; from my own world of scholarly publishing). There's clearly a lot of work to be done. And the reality is that copyrights and patents and trademarks and norms and software and hardware are going to be hard to reconcile into a simple, single license that "makes copyleft hardware" a reality. But it was fun to be in a room with so many passionate, brilliant people who want to make the world a better place through collaborative research.&lt;/p&gt;

&lt;p&gt;More to come once results emerge...&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/03/open_hardware.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/kTCjhVX0N9Y" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/03/open_hardware.php</guid>
         <category />
         
         <pubDate>Mon, 22 Mar 2010 19:09:52 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/03/open_hardware.php</feedburner:origLink></item>
      
      <item>
         <title>Science Commons T-shirt Contest</title>
          <description>&lt;p&gt;&lt;a href="http://scienceblogs.com/commonknowledge/assets_c/2010/02/scienceCommonsTeeNavy-41369.php" onclick="window.open('http://scienceblogs.com/commonknowledge/assets_c/2010/02/scienceCommonsTeeNavy-41369.php','popup','width=500,height=577,scrollbars=no,resizable=no,toolbar=no,directories=no,location=no,menubar=no,status=no,left=0,top=0'); return false"&gt;&lt;img src="http://scienceblogs.com/commonknowledge/assets_c/2010/02/scienceCommonsTeeNavy-thumb-500x577-41369.jpg" width="500" height="577" alt="scienceCommonsTeeNavy.jpg" class="mt-image-none" style="" /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The winner in our t-shirt contest. The background to the robot is a 2d barcode that de-references back to creativecommons.org, which I like a lot :-)&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/science_commons_t-shirt_contes.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/QpIcl5dyYcE" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/science_commons_t-shirt_contes.php</guid>
         <category />
         
         <pubDate>Sun, 21 Feb 2010 06:41:51 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/science_commons_t-shirt_contes.php</feedburner:origLink></item>
      
      <item>
         <title>Reaching Agreement On The Public Domain For Science</title>
          <description>&lt;p&gt;&lt;a href="http://scienceblogs.com/commonknowledge/assets_c/2010/02/4370186974_363d182500-41316.php" onclick="window.open('http://scienceblogs.com/commonknowledge/assets_c/2010/02/4370186974_363d182500-41316.php','popup','width=500,height=384,scrollbars=no,resizable=no,toolbar=no,directories=no,location=no,menubar=no,status=no,left=0,top=0'); return false"&gt;&lt;img src="http://scienceblogs.com/commonknowledge/assets_c/2010/02/4370186974_363d182500-thumb-500x384-41316.jpg" width="500" height="384" alt="4370186974_363d182500.jpg" class="mt-image-center" style="text-align: center; display: block; margin: 0 auto 20px;" /&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo outside the Panton Arms pub in Cambridge, UK, licensed to the public under Creative Commons Attribution-ShareAlike by &lt;a href="http://www.flickr.com/photos/jwyg/"&gt;jwyg&lt;/a&gt; (Jonathan Gray).&lt;/p&gt;

&lt;p&gt;Today marked the public announcement of a set of principles on how to treat data, from a legal context, in the sciences. Called the &lt;a href="http://pantonprinciples.org/"&gt;Panton Principles&lt;/a&gt;, they were negotiated over the summer between myself, &lt;a href="http://www.rufuspollock.org/about/"&gt;Rufus Pollock&lt;/a&gt;, &lt;a href="http://blog.openwetware.org/scienceintheopen/"&gt;Cameron Neylon&lt;/a&gt;, and &lt;a href="http://en.wikipedia.org/wiki/Peter_Murray-Rust"&gt;Peter Murray-Rust&lt;/a&gt;. If you're too busy to read them directly, here's the gist: publicly funded science data should be in the public domain, full stop.&lt;/p&gt;

&lt;p&gt;If you know me and my work, this is nothing new. We have been &lt;a href="http://creativecommons.org/weblog/entry/7917"&gt;saying this&lt;/a&gt; since late 2007. I've already gotten a dozen emails asking me why this is newsworthy, when it's actually a less normative version of the &lt;a href="http://sciencecommons.org/projects/publishing/open-access-data-protocol/"&gt;Science Commons protocol for open access to data&lt;/a&gt; (we used words like "must" and "must not" instead of the "should" and "should not" of the principles).&lt;/p&gt;

&lt;p&gt;It's newsworthy to me because it represents a ratification of the ideals embodied in the protocol by two key groups of stakeholders. First, real scientists - Cameron and Peter are two of the most important working scientists in the open science movement. Getting real scientists into the fold, endorsing the importance of the public domain, is essential. They're also working in the UK, which has some copyright issues around data that can complicate things in a way we forget about here in the post-colonial Americas.&lt;/p&gt;

&lt;p&gt;Second, it's newsworthy because Rufus and I both signed it. Rufus helped to start the &lt;a href="http://okfn.org/"&gt;Open Knowledge Foundation&lt;/a&gt;, and he's an important scholar of the public domain. We're in many ways in the same fraternity - we care about "open" deeply, and we want the commons to scale and grow, because we believe in its role in innovation and creation...indeed, in its role in humanity. &lt;/p&gt;

&lt;p&gt;But we're on different sides of a passionate debate about data and licenses. I'm not going to recapitulate it here, you can find it in the googles if you want. Suffice to say we have argued about the role of the public domain as a first principle in general for data, as opposed to the specifics of data in public funded science. But for both of us to sign onto something like this means that even in the midst of heated argument we can find common ground - public money should mean public science, no licenses, no controls on innovation and reuse, globally.&lt;/p&gt;

&lt;p&gt;It's important for the science part. It's also a good lesson, I hope, that even those of us who find themselves on opposite sides of arguments inside open are usually fighting for the same overall goals. I'll keep arguing for my points, and Rufus will keep arguing for his, but that should never keep us from remembering the truly common goals we share inside the movement. I'm proud to be a part of it.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/reaching_agreement_on_the_publ.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/14Aje1vgKHY" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/reaching_agreement_on_the_publ.php</guid>
         <category />
         
         <pubDate>Fri, 19 Feb 2010 12:24:42 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/reaching_agreement_on_the_publ.php</feedburner:origLink></item>
      
      <item>
         <title>Tech4Society, Day 3</title>
          <description>&lt;p&gt;I'll start my final post on the Tech4Society conference by giving thanks to the &lt;a href="http://ashoka.org/"&gt;Ashoka&lt;/a&gt; folks for getting me here to be a part of this conference. Most of the time, even in the developing world, I'm surrounded by &lt;a href="http://en.wikipedia.org/wiki/Digital_native"&gt;digital natives&lt;/a&gt;, or people who emigrated to the digital nation. It's an enveloping culture, one that can skew the perception of the world to one where everyone worries about things like copyrights and licenses, and whether or not data should be licensed or in the public domain. &lt;/p&gt;

&lt;p&gt;There's a big world of entrepreneurs out there just hacking in the real world. First life, if you will. High touch, not high tech. &lt;/p&gt;

&lt;p&gt;Being enveloped in their world for a few days gave me a lot of new perspectives on the open access and open educational resources movements. As always with this blog, my intention to write may exceed my delivery of text, but I'm going to try to chew through the perspectives. Getting off the road in a few weeks is going to help. &lt;/p&gt;

&lt;p&gt;But I now get at a deep level the way that obsessive cultures of information control in the scholarly and educational literature represent a high tax, inbound and outbound, on the entrepreneur, whether social or regular. If you don't know the canon, you're doomed to repeat it. And we don't have the time, the money, or the carbon to repeat experiments we know won't work. We can't afford to let good ideas go un-amplified, because we need tens of thousands of good ideas.&lt;/p&gt;

&lt;p&gt;At my panel today on scale, we focused mainly on why scale is hard, the problems of scale. The CC experience - going from 2 people in a basement at Stanford to 50 countries in 6 years - is an example of what I called "catastrophic success". It's a nice way to think of what I also like to call the Jaws moment, after the scene in the 1970s action film where, having hoped to find a shark to catch, they find one muuuuuch bigger than they expected. The relevant quote is "&lt;a href="http://www.youtube.com/watch?v=8gciFoEbOA8"&gt;we're gonna need a bigger boat"&lt;/a&gt; - and that is what happens sometimes at internet scale. Entrepreneurs need to know why they want to scale, what scale means to them, and how to measure success, especially social entrepreneurs. Because if cash isn't the only metric, the metrics you choose will wind up defining your success at scale.&lt;/p&gt;

&lt;p&gt;There was a great question about scaling passion. I am going to try and address that in another post. I'm not quite in a mental state to get that post out yet, though.&lt;/p&gt;

&lt;p&gt;It wasn't just the social entrepreneurs, but also CC community experiences. &lt;a href="http://www.ted.com/profiles/view/id/299148"&gt;Gautam John &lt;/a&gt; challenged me, eloquently and at length, about the way that Creative Commons engages with its community. I went into the argument convinced of my position, and left much less so. That's as good as arguments get for me. &lt;/p&gt;

&lt;p&gt;Ashoka and Lemelson foundations are doing great work, supporting inventors around the world (though I would have liked to have seen some Eastern bloc inventors - a curious lack of Slavic accents - wonder why). It was an honor to crash their party.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_3.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/4ShQEV0MRkg" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_3.php</guid>
         <category />
         
         <pubDate>Sat, 13 Feb 2010 14:02:32 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_3.php</feedburner:origLink></item>
      
      <item>
         <title>Tech4Society, Day 2</title>
          <description>&lt;p&gt;Getting ready to head up to &lt;a href="http://tech.ashoka.org/hyderabad_info"&gt;Tech4Society's final day&lt;/a&gt;. I'm on a panel called the tipping point, about how to scale social entrepreneurial success beyond a local region or state. My instinct is to say "pack your suitcase and start traveling" but that's not very helpful. Even if it's how I have been approaching the problem.&lt;/p&gt;

&lt;p&gt;Yesterday I wasn't on a panel. It was a good moment to do some listening. I sat in on a few panels, but was most moved by the trends in Africa session. In other trends panels, the trends were things like "open source" - positive trends. In Africa it was all about how difficult the governance problems are, how an innovator or social entrepreneur is looked on with at best skepticism or out worst outright hostility, by both local society and by the government. &lt;/p&gt;

&lt;p&gt;It was still amazing to hear the breadth of ingenuity at work. I heard about training rats to sniff out landmines, clay refrigerators that allow girls to go to school rather than hawking the harvest before it spoils...and in the same breath, about how it takes five hours to get one hour of work done, because of the difficulty of keeping a steady power supply.&lt;/p&gt;

&lt;p&gt;At lunch I crashed the Indonesian table, where I was asked if I was part of the youth venture group. Nicest age-related compliment I've gotten in a while (the youth venture folks are like 16 years old). But it does strip away any pretense of gravitas I thought I might have had. &lt;/p&gt;

&lt;p&gt;I also got to spend some quality time with Richard Jefferson of &lt;a href="http://www.cambia.org/"&gt;CAMBIA&lt;/a&gt;. Richard is a seasoned social entrepreneur who has been hacking away at the patent problem in "open" biotech for about 20 years now. I always learn a lot from him.&lt;/p&gt;

&lt;p&gt;At the end of the day the heat and the jetlag caught me, and I fell asleep before the dinner, which is a bummer. &lt;/p&gt;

&lt;p&gt;I'm looking forward to having some time off the road in a few weeks to try and integrate this experience with the other travel over the past four months. There's a long way between the World Economic Forum at Davos and this. The entrepreneurs here are doing what they do against such long odds that it can make the whole "cult of the successful entrepreneur" in the US look kind of lame. &lt;/p&gt;

&lt;p&gt;It doesn't take a hero to make a social networking site, it just takes some Ruby code. We have layer upon layer upon layer of infrastructure that makes it easy to innovate in the US. We have stable power grids, for the most part, and communications lines. You can buy a computer for under $500, slap Linux on it, and you're ready to start a software company. You don't have to pay a registration fee that takes six months, or worry that the government is going to crack down on you (despite what some crackpots may think) if you protest or run a business that disagrees with the ruling elites. &lt;/p&gt;

&lt;p&gt;That level of social, political, and technical infrastructure lifts us all up who benefit from it. It's invisible to most of us most of the time, and it's a good thing to be reminded that it's not something to be taken for granted.&lt;/p&gt;

&lt;p&gt;Off to day 3.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_2.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/iRSLvvDPK3Q" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_2.php</guid>
         <category />
         
         <pubDate>Fri, 12 Feb 2010 21:31:34 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/tech4society_day_2.php</feedburner:origLink></item>
      
      <item>
         <title>On the Nature of Ideas</title>
          <description>&lt;p&gt;I did an interview recently where the author, clearly having done some homework, called out an old quote of mine arguing that ideas aren't like widgets or screws, that they're not industrial objets.&lt;/p&gt;

&lt;p&gt;I'd said that a long time ago, inspired by John Perry Barlow's&lt;a href="http://"&gt; Declaration of Independence of Cyberspace&lt;/a&gt;. Here's the money quote: "Your increasingly obsolete information industries would perpetuate themselves by proposing laws, in America and elsewhere, that claim to own speech itself throughout the world. These laws would declare ideas to be another industrial product, no more noble than pig iron. In our world, whatever the human mind may create can be reproduced and distributed infinitely at no cost. The global conveyance of thought no longer requires your factories to accomplish."&lt;/p&gt;

&lt;p&gt;The world that John Perry was talking about has not come to pass, completely. The governments have certainly moved to impose more and greater controls. But as Lessig noted just a few years later in &lt;a href="http://www.code-is-law.org/"&gt;Code: and Other Laws of Cyberspace&lt;/a&gt;, the aspects of cyberspace that promised liberation, a nation of the Mind...those aspects were the output of human-controlled systems, and humans could and would change the rules if they didn't like the outcomes. &lt;/p&gt;

&lt;p&gt;I was there for parts of these conversations. Gave JPB a ride around town, harassing him about the declaration and about Cassidy. I put together Lessig's book party for Code when it came out. But the thing about ideas stuck with me more than the rest. &lt;/p&gt;

&lt;p&gt;I'd studied epistemology, the theory of knowledge. You get a lot of examples of attempting to codify ideas (and brains, the storage tanks of ideas) into the machinery of the time (See the masterful book, "&lt;a href="http://www.amazon.com/Memory-Practices-Sciences-Inside-Technology/dp/0262524899/ref=sr_1_1?ie=UTF8&amp;s=books&amp;qid=1265925842&amp;sr=1-1"&gt;Memory Practices in the Sciences&lt;/a&gt;" for more). But in the end, ideas resist complete capture.&lt;/p&gt;

&lt;p&gt;They're ethereal. We've spent thousands of years trying to codify them, into Plato's forms, into machines, now into code. The dominant industrial paradigm tends to be the stuff we use to try and understand them and their human substrates - pumps and machines to explain the brain in the industrial age, circuits and pathways in the digital. This ethereal nature makes it hard to get the ideas into the powerful information systems of the day, which are based on bits and bytes. It's one of the reasons that the most powerful idea transmissions systems we have are humanist - text, sound, video. It's why something as lousy as powerpoint can take over, because it's a way for people to talk to people.&lt;/p&gt;

&lt;p&gt;It's hard to make ideas into widgets or screws because of this. It's also hard because we all see the world differently, even those of us who agree. We use common words as proxies to help convey that this red ball is an apple and this green ball is also an apple. Making the word apple into an abstracted computation tool is hard, because you have to decide what it means, and convince others to use your meaning rather than their own. &lt;a href="http://en.wikipedia.org/wiki/Cyc"&gt;Cyc's been pushing on this for 25 years&lt;/a&gt; and we still don't have the Star Trek computer recognizing our voices.&lt;/p&gt;

&lt;p&gt;But we're starting to have to try to make ideas at least representable as widgets. The problem is that the information space is overwhelming us as people. We can, using robots in the lab, sensor networks in the ocean, miniature microphones in public spaces, genotype smears on red light signals, generate data at such a level that we simply cannot use our own brains to proces the data into an information state that lets us extract, test, and generate ideas.&lt;/p&gt;

&lt;p&gt;There's two things we can do, one easy and one hard. First, we can make the existing technologies for idea transmission (writing it down onto paper and publishing it) more democratic and network friendly. That starts with good formats: putting ideas into PDFs is a terrible idea. The format blocks the ability to take the text out, remix it, translate it, reformat it, text mine it, have it read to a blind person via text-to-speech, and on and on. It continues with open access (so we don't create a digital divide first, and so we enable the entrepreneurs of the world, wherever they are). &lt;/p&gt;

&lt;p&gt;I'm at a conference in Hyderabad called &lt;a href="http://tech.ashoka.org/hyderabad_info"&gt;Tech4Society&lt;/a&gt; that is packed to the gills with inventors and social entrepreneurs, who for the most part have no access to the scientific and technical literature. It's all in English - which many, but not all, speak here. It's very expensive - nuclear physics  journals can cost more - per year - than a new car. And this is a tax on the entrepreneurs of the world. &lt;/p&gt;

&lt;p&gt;Inventors have to invent. It's in their blood. And they have the capacity to rapidly combine information from multiple sources to assemble new projects. I heard today of systems that leverage sugar palms in Indonesia to power villages, of local decentralized power panels for wind and solar to give each house its own power, and more and more and more. But this is being done without the newest knowledge, knowledge that is on the web somewhere...but locked up by paywalls. &lt;/p&gt;

&lt;p&gt;We as Americans send a lot of money. We'd be a damned sight better if we sent a lot of knowledge. &lt;/p&gt;

&lt;p&gt;The Open Access movement is being driven mainly inside the developed world. US and EU librarians feel the pinch of the serials pricing crisis, and funders like the US National Institute of Health and the Wellcome Trust take policy directions that lead towards the availability of the biomedical research. And it's wonderful that the solutions to these problems all lift the developing world along the way. It seems that the scholarly literature will, in fits and starts, and in some disciplines faster or slower, find its proper place on the net, free of commercial restrictions, one of these days.&lt;/p&gt;

&lt;p&gt;But it's not just ideas, it's what to do with the ideas. Richard Jefferson today made the lovely point that the patent literature is a giant database of recipes to make inventions. And that if you can find the inventions that were patented in the US, but not in India, you've got a lot of good stuff to work on in India. This is true. And deeply important. &lt;/p&gt;

&lt;p&gt;But I got a little melancholy thinking of the stuff that comes before an inventor becomes a social entrepreneur, ready to apply for funding or speak in front of 200 people at a conference. Maybe they can't read the patents and understand the information. Maybe they just need to build some furniture for their house, or fix the stove. I had a sense-memory of long shelves of the books in Home Depot, the how-to guides, the recipes for doing simple stuff, unpatented stuff, but essential stuff, and I look at the amazing user-driven innovative spirit that rules the day in India, and I want to cry at the amount of knowledge that is deprived. Give these folks the books and get out of the way!&lt;/p&gt;

&lt;p&gt;I wish we could come together as a culture and create an open source set of how-to books to parallel the scholarly literature. Those book are how I learned to rewire sockets, to fix plumbing. Where I learned what was dangerous and what was safe.  They're a place where those ideas, laid out in the papers that are becoming free, became methods that I could use. Where the ideas became actionable for me. Imagine if those books were movable from my server where I wrote them, to a server in Africa who translated them into Kiswahili, or Chichewa. If they could be formatted to be read on the mobile phones ubiquitous across the world. If they could lead to one more hour of light per night through the creation of lightweight photovoltaics. &lt;/p&gt;

&lt;p&gt;Has anyone out there done this yet? Anyone interested in doing it? Anyone immediately get a rash and freak out? All of those reactions are interesting to me. &lt;/p&gt;

&lt;p&gt;The second part of why ideas are hard will have to wait for the next post. Suffice to say the word "semantic" will feature prominently.&lt;/p&gt;

&lt;p&gt;I'll post more from day 2 tomorrow. Jetlag over and out.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/on_the_nature_of_ideas.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/m7Nw8qom2Cc" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/on_the_nature_of_ideas.php</guid>
         <category />
         
         <pubDate>Thu, 11 Feb 2010 16:35:44 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/on_the_nature_of_ideas.php</feedburner:origLink></item>
      
      <item>
         <title>Quick Roundup</title>
          <description>&lt;p&gt;I'm going to get re-starting blogging here this week, after a crazy few weeks of work at the day job.&lt;/p&gt;

&lt;p&gt;In the meantime, here's a few links to stuff I'm doing, reading, and writing.&lt;/p&gt;

&lt;p&gt;The team has been busy working on governance and systems for the &lt;a href="http://www.sagebase.org/COMMONS/Repository.html"&gt;Sage Commons&lt;/a&gt; (the first globally coherent dataset is available - you have to register, but it's open data...). &lt;/p&gt;

&lt;p&gt;We've also been busy pushing on the first beta version of the Creative Commons &lt;a href="http://sciencecommons.org/projects/patent-licenses"&gt;patent tools &lt;/a&gt; for the &lt;a href="http://greenxchange.force.com"&gt;GreenXchange&lt;/a&gt; launch in Davos. There is more documentation on GreenXchange in the &lt;a href="http://sciencecommons.org/resources/readingroom/"&gt;reading room over at SC&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;I contributed an essay to the &lt;a href="http://research.microsoft.com/en-us/collaboration/fourthparadigm/"&gt;Microsoft Research book on the Fourth Paradigm&lt;/a&gt;, which is dedicated to Jim Gray. I never got to meet Jim - we were scheduled to chat the Wednesday after his boat disappeared - but I'm honored to be part of this collection. &lt;/p&gt;

&lt;p&gt;The folks at &lt;a href="https://open.umich.edu/blog/2010/01/29/geographical-imagery-and-data-of-michigan-available-for-all/"&gt;Michigan View have released a bunch of amazing Landsat data under CC0&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;Michigan continues to be a leader in thinking and doing open data - you should also check out the &lt;a href="https://cabig-kc.nci.nih.gov/DSIC/KC/index.php/Main_Page "&gt;Data Sharing and Intellectual Capital workspace&lt;/a&gt; work going on there, which is part of the National Cancer Institute's broader caBIG initiative.&lt;/p&gt;

&lt;p&gt;Hope Leman's &lt;a href="http://significantscience.com"&gt;Significant Science&lt;/a&gt; blog is becoming a go-to place for open science. She's done a magnificent &lt;a href="http://significantscience.com/2010/01/28/the-indispensable-man-of-open-science-a-talk-with-cameron-neylon/"&gt;interview with Cameron Neyl&lt;/a&gt;on. As Donna Wentworth noted, it's a great enthographic overview of open science. &lt;/p&gt;

&lt;p&gt;We're doing an &lt;a href="http://scs.eventbrite.com/"&gt;event&lt;/a&gt; that Hope is kindly helping out with in Seattle in a few weeks - if you're in the Pacific Northwest, come on by. And if you have any ideas for the &lt;a href="http://sciencecommons.org/weblog/archives/2010/01/27/design-a-new-t-shirt-for-science-commons-and-win-a-trip-to-seattle-to-attend-science-commons-symposium---pacific-northwest/"&gt;next t-shirt design for us&lt;/a&gt;, you could win a ticket...&lt;/p&gt;

&lt;p&gt;And that'll do for now. I'll be back with some semi-coherent posts soon, I hope.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2010/02/quick_roundup.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/Msp6etkJIIE" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2010/02/quick_roundup.php</guid>
         <category />
         
         <pubDate>Mon, 01 Feb 2010 11:07:59 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2010/02/quick_roundup.php</feedburner:origLink></item>
      
      <item>
         <title>In Which We Continue To Push The Sisyphean Rock Up The Hill</title>
          <description>&lt;p&gt;My &lt;a href="http://scienceblogs.com/commonknowledge/2009/10/open_source_science_or_distrib.php"&gt;last&lt;/a&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/11/distributed_science_part_2.php"&gt;posts&lt;/a&gt; on why I don't like the open source metaphor for science have generated a lot of good comments, here and in my email, twitter, and in person. &lt;/p&gt;

&lt;p&gt;They've forced me to think about what exactly it is about the meme that makes me so uncomfortable, and raised some good objections and points. I'm going to try to chew through a few of them in this post and then ditch the topic for a while, as I've got a lot of complaining to do about publishing and data and those topics have had to take a back seat for a few weeks while I worked this through my system. &lt;/p&gt;

&lt;p&gt;On a side note, I actually kinda felt like a real blogger the last few weeks. &lt;/p&gt;

&lt;p&gt;I guess for me the open source metaphor is so tied to software that its applicability as a metaphor is limited. I have done a very informal, personal, anecdotal, but multi-year, survey of people I talk to on this topic. For most people "open source" is an idea seen through a glass darkly, a vague mishmash of ideas of political freedom, distributed development methodology, and magical legal tools. &lt;/p&gt;

&lt;p&gt;"We need open source [insert variable] to do [insert task currently performed by big evil company]" is almost an algorithm of faith in my world. I hear it again and again, marked by almost no understanding of the context in which open source &lt;strong&gt;&lt;em&gt;software&lt;/em&gt;&lt;/strong&gt; actually exists and operates. This is what I'm on about. Open source isn't a magic incantation we can use to summon a community and create a public good.&lt;/p&gt;

&lt;p&gt;Open "source" to most who use the metaphor is so much more than "the source is available" that we need to do some pushing back against it, even as we must also celebrate the intentions behind its use. &lt;/p&gt;

&lt;p&gt;I'm going to make a few attempts to untangle the mishmash. &lt;/p&gt;

&lt;p&gt;First, open source came from Free Software. If you haven't read through the &lt;a href="http://en.wikipedia.org/wiki/Free_and_open_source_software"&gt;histories of Free v. Open&lt;/a&gt;, please go do so.  But I would loosely generalize that Free is more about Freedom, of programmers, of speech, and of society, whereas Open Source is more about a development methodology embracing distribution of tasks and interconnectivity of outputs. They have a lot in common, but they're not the same.&lt;/p&gt;

&lt;p&gt;Second, both Free and Open Source depend upon a public approach to copyright, which is the &lt;a href="http://en.wikipedia.org/wiki/Open_source_license"&gt;open copyright license&lt;/a&gt;. The existence of a powerful, relatively internationally harmonized property right is absolutely essential to the entire open source enterprise. Another key point in copyright is that the creator-programmer owns all her rights necessary to license those rights (absent signing them away to a company or other institution in a contract, of course). If she writes code, she owns it, without applying to a central authority, for a hell of a long time. &lt;/p&gt;

&lt;p&gt;This power is at the root of the power of the license. It cannot be understated, and I'll come back to it later, because the absence of such a right that works this way is a central flaw to the naive application of the metaphor in science. &lt;/p&gt;

&lt;p&gt;Third, open source software hasn't changed the world just because it was free, or openly licensed. It sits on top of an infrastructure that was highly leveraged to support something like open source - the &lt;a href="http://en.wikipedia.org/wiki/Internet_Protocol_Suite"&gt;internet stack&lt;/a&gt;, the explosion of microcomputers, the magic intersection of moore's and metcalfe's and joy's laws, the democratization of network access, and more. And on top of all of this was also the explosion in programming tools, object orientation, and modularity of software design. &lt;/p&gt;

&lt;p&gt;Let me phrase it as a question. Would the four freedoms and the GNU GPL have been sufficient to create an explosion of free and open source software in the mid-60s? Pre-internet, pre-web, in the days of mainframes and timesharing and tiny memory and machine code? &lt;/p&gt;

&lt;p&gt;I would propose the answer is no.&lt;/p&gt;

&lt;p&gt;These three elements are poorly represented in science. We have some desire for the first issue - starting with Freedom. That's probably the most advanced. And that's why the open source science movement starts with appropriation of language and metaphor from software. I understand it. I support the ideas behind it. However I think it blinds us to the things that block the intentions from being realized, which are many. I'll expand on two here.&lt;/p&gt;

&lt;p&gt;First, the legal basis for open licensing in science is not simple, powerful, and internationally harmonized. &lt;/p&gt;

&lt;p&gt;Science creates at least four classes of knowledge artifacts at their most basic level: creative works (whether in a journal or a webby form like a blog, whether narrative or photo or video), data (whether "raw" or processed), databases (which are different from data, and may contain creative works as well as data), and inventions (which may or may not be patented). Each of these four classes carries its own often wacky property rights regimes, some of which are amenable to open source style licenses, some of which aren't, and some of which it's utterly unclear if open source will work or not. &lt;/p&gt;

&lt;p&gt;To make things worse, science takes place in institutions. That means institutional claims on property rights. Institutions have offices set up specifically to exploit property rights, not share them. And even if you can get the institution on board, the creator usually does not own all the rights necessary to make the kind of freedom available - remember, this all starts with Freedom - as we contemplate in the open source metaphor. Worse yet again, getting the property right associated with inventions (patent) costs a ton of money - so giving it away as soon as you get it is much harder as a value proposition than in copyright, which descends from the heavens when the pen lifts from the paper.&lt;/p&gt;

&lt;p&gt;Patents and copyrights don't mix beautifully, either. If I own a copyright on a gel box design, I can release the design and "make the source code available" - but if my neighbor owns a patent on it, that neighbor can sue anyone who tries to actually build the gel box. This is something of a problem in software. But it's a massive problem in science. Especially life sciences, which are built on patents as proxies for economic value and create enormous employment opportunities for attorneys as a result.&lt;/p&gt;

&lt;p&gt;Data and databases are another place where the underlying property regimes don't work as well for open source as in software. But that's difficult enough to merit its own post. Suffice to say if Open Data had a facebook page, its relationship status with the law would be "It's Complicated."&lt;/p&gt;

&lt;p&gt;The second block is the insufficient mix of infrastructure. We can make creative works available, we can post data, we can license (maybe) inventions, we can integrate databases. But stitching it all together is the hard part. It is hard to compile four classes of knowledge products, much harder than compiling software. And the open source metaphor again builds the expectation that if we "make it open" that we'll get a magic network effect, that wikipedia will emerge for science. &lt;/p&gt;

&lt;p&gt;But the infrastructure for software isn't strong enough to stitch together science knowledge. Most science knowledge is locked up in PDF and Word formats, lacks hyperlinks, or in standalone databases. It's not "modular" in the sense that software is, even though it's just as socially constructed as software in its own way. We've designed science knowledge for a human operating system, not a computerized one. &lt;/p&gt;

&lt;p&gt;This is why I'm semi-obsessed with building linked data infrastructure, semantics, ontologies, and so on. It is going to be essential to realizing the intentions behind the open source metaphor - that knowledge connected becomes more valuable than the sum of its parts, that many of us can work separately on the same task and create a common good. We're in the pre-internet world of science, metaphorically, and we need to build the networks and the protocols first, we need the machines to get cheaper and ubiquitous, we need common languages for data and concepts - then we can start talking about a free software metaphor being accurate.&lt;/p&gt;

&lt;p&gt;I'm not beating up on the metaphor because I hate the idea. And if I can find evidence that people are using the metaphor in full understanding of the realities between here in science, and there in open source science, I'll dial it back. &lt;/p&gt;

&lt;p&gt;But so far I haven't found that evidence. And I think propagating the open source metaphor - without a hard-eyed examination of the barn raising we have to do before we get anything as transformative for science as GNU/Linux has been for software - risks hiding the hard stuff and creating unrealistic expectations that could boomerang on us all.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/12/in_which_we_continue_to_push_t.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/yJi3SWoaHfQ" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/12/in_which_we_continue_to_push_t.php</guid>
         <category />
         
         <pubDate>Thu, 03 Dec 2009 10:02:58 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/12/in_which_we_continue_to_push_t.php</feedburner:origLink></item>
      
      <item>
         <title>Distributed Science, Part 2</title>
          <description>&lt;p&gt;I got a lot of feedback on my&lt;a href="http://scienceblogs.com/commonknowledge/2009/10/open_source_science_or_distrib.php#comments"&gt; last post&lt;/a&gt; in which I argued that open source is the wrong metaphor fo science, because it ties us too closely to the artifact that is open source software. The core of my argument remains the same - &lt;em&gt;science is not software, and we shouldn't treat it the way we treat software&lt;/em&gt;. But I got a few comments, here on the blog and in email, that are worth looking at.&lt;/p&gt;

&lt;p&gt;Here's comment #1.&lt;/p&gt;

&lt;blockquote&gt;You cite openwetware and the biobricks registry, but if you look closer, openwetware is a wiki, not a website about open source wetware tech. To my knowledge, other than the people over at diybio, there have been no signs of anyone with an understanding of free and open source software infrastructure (not the legalese- the toolchains) applying the concepts to the world of open source science.&lt;/blockquote&gt;

&lt;p&gt;This comment illustrates my point by missing it, which is that we should not be applying the understanding of software to science. In software, we the humans are in charge. We write the code. We compile it. Everything exists inside a system that we built, that is at least somewhat intelligently designed. Bringing this "understanding" to science means we shove a science peg into a software slot. The idea that "open source science" should be a site about wetware tech betrays a focus on the construction of tech, which is indeed the point of software.&lt;/p&gt;

&lt;p&gt;But science isn't like software. Science is about extending the boundaries of our ignorance, not making technology. The difference between making technology (which is the point of software) and making discoveries (the point of science) is the root of the failure of the "open source science" metaphor. Science is about creating knowledge that doesn't exist and exposing ignorance that does exist, not about writing source code that we control. &lt;/p&gt;

&lt;p&gt;In honor of his recent passing, here's Claude Lévi-Strauss: "The scientist is not a person who gives the right answers, he's one who asks the right questions." (from&lt;a href="http://fr.wikipedia.org/wiki/Le_Cru_et_le_cuit"&gt; Le Cru et le cuit&lt;/a&gt;, 1964)&lt;/p&gt;

&lt;p&gt;This is precisely why I want to take us up a layer in the ontology. Open source software is an example of distributed innovation, and as an inspiration to make distributed innovation happen in science, it's lovely. But it's an inspiration, not a map.&lt;/p&gt;

&lt;p&gt;We should absolutely have distributed innovation in science. Open WetWare (which I am well aware is a wiki) contains many protocols, crafts and techniques, that are shared openly. This is a locally relevant form of distribution, even if it doesn't fit into an open source software box. Control over protocols and craft is at the core of one of the biggest resistors to distribution in science, which is &lt;a href="http://www3.interscience.wiley.com/cgi-bin/fulltext/121633537/HTMLSTART?CRETRY=1&amp;SRETRY=0"&gt;competitive withholding&lt;/a&gt;. So is the registry of standard biological parts. These are resources and toolchains that absolutely support distribution of capability and increase capacity, which are fundamental to early-stage distributed innovation. &lt;/p&gt;

&lt;p&gt;They're just not what we expect when we wear open source glasses.&lt;/p&gt;

&lt;p&gt;Here's comment #2:&lt;/p&gt;

&lt;blockquote&gt;The "Open Gel Box" project is an initiative to bring biotech equipment into the 21st century. We need innovation in "established" tools to make them intuitive and accessible for anyone who wants to work with DNA. To that end, a group of users from the DIYbio list got together and designed a better, faster gel system than what exists today.

&lt;p&gt;Pearl Biotech is now manufacturing a complete gel electrophoresis system according to the Open Gel Box design The Pearl Gel Box is available for $199 at http://www.pearlbiotech.com. We're advocating for better equipment on all fronts, such as an Open Thermal Cycler.&lt;/blockquote&gt;&lt;/p&gt;

&lt;p&gt;I think this is awesome. It's not "open source" though. It's not even what I'd call "distributed innovation" - the innovation theorists call this kind of thing &lt;a href="http://en.wikipedia.org/wiki/User_innovation"&gt;User-Driven Innovation&lt;/a&gt;. This is about as clear a case of UDI as I know, right down to the fact that it's designed by the DIY folks and then made pretty and sold by a company. This again gets to the paucity of the open source software example. It simply isn't big enough to fit science into it. &lt;/p&gt;

&lt;p&gt;Distributed science, user-driven science, open innovation science, we need ALL of them, not a narrow idea that comes from software. It's about hardware for science. It's about data for science. It's about laboratories for science. It's about research departments and funders and promotion and tenure. It's about paradigms, and paradigm shifts. &lt;/p&gt;

&lt;p&gt;It's not software. &lt;/p&gt;

&lt;p&gt;We control software. We don't control science. DIY Biology is one of the absolute leading examples of how, when we have a critical mass of open craft and protocols, users can lead the way. But it's not something that's enabled by an open source license, a code version repository, and other hallmarks of open source software. It's users saying, "screw this, I can do better" - and doing it. It's users who know the problem best and design the best solutions. &lt;/p&gt;

&lt;p&gt;The business school folks call this "&lt;a href="http://web.mit.edu/evhippel/www/papers/stickyinfo.pdf"&gt;stickiness&lt;/a&gt;." The knowledge of how to make the solution is localized - sticks - to the user. The dumb firms in the sector only make products their marketing departments tell them about, and the smart ones find ways to take user inventions and turn them into their product lines. Like Pearl. &lt;/p&gt;

&lt;p&gt;Comment #3:&lt;/p&gt;

&lt;blockquote&gt; (from my post: Stem cells, mice, vectors, plasmids, and more will need to available outside the old boy's club that dominates modern life sciences.)

&lt;p&gt;This is simply never gonna happen, because of the huge irreducible expense of maintaining and manipulating these reagents.&lt;/blockquote&gt;&lt;/p&gt;

&lt;p&gt;See: &lt;a href="http://personalgenome.org"&gt;Personal Genome Project&lt;/a&gt;, &lt;a href="http://ccr.coriell.org/"&gt;Coriell Cell Culture Repository&lt;/a&gt;, &lt;a href="http://www.jax.org"&gt;Jackson Laboratories&lt;/a&gt;, &lt;a href="http://www.straininfo.ugent.be/"&gt;StrainInfo&lt;/a&gt;. I could link a dozen more. The nodes are emerging. What's missing is the network that connects them. What's missing is an impact factor for materials. &lt;/p&gt;

&lt;p&gt;We're headed straight towards a future where scientists will need to publish their tools, data, and narratives, instead of compressing everything into a "paper" that is constrained by the cost of printing and mailing. I for one can't wait. It's going to be a key to distributing democratized access to tools, which is fundamental for both distributed innovation *and* user-driven innovation.&lt;/p&gt;

&lt;p&gt;Comment #4:&lt;/p&gt;

&lt;blockquote&gt;I believe your historical facts are a little skewed. Open Biology perhaps began on the internet back with BIONET, which functioned well through the late 80's and early 90's, until the network apparently failed to grab sufficient interest for funding. [...] There have been efforts to create biology software repositories (similar to sourceforge.net except for Biology software) and these have largely failed to attract a majority of Bio-scientists too.&lt;/blockquote&gt;

&lt;p&gt;This comment's talking about software. I'm not. It again illustrates the way that the open source metaphor comes with code-centric blinders. &lt;/p&gt;

&lt;blockquote&gt;It would be great to accelerate this process even further, for example by expanding PLoS, encouraging all scientists to publish their working software (for example, MATLAB scripts) into open source repositories&lt;/blockquote&gt;

&lt;p&gt;Now this is talking about the foundations for distributed science. When there is software in science, it should be published. Just like stem cells. Into repositories. Couldn't agree more.&lt;/p&gt;

&lt;blockquote&gt;encouraging the people-in-the-middle (hobbyists, engineers) to publish in an intermediate form which isn't as strict as a scientific journal yet maintains some level of technological standard and legitimacy -- similar to the Internet RFC's, which started as simple technical memo's.&lt;/blockquote&gt;

&lt;p&gt;Now here's where the comment truly shines, IMO. This is thinking broadly about breaking open the central metaphor of knowledge governance in science. This is not about "open source" - the internet RFCs aren't "open source software" - they are protocols, distributed for implementation and comment. Sort of like that stuff on the Open WetWare wiki, huh?&lt;/p&gt;

&lt;p&gt;Coming back to my point. &lt;/p&gt;

&lt;p&gt;Let's take off the open source glasses. Making science isn't like making software. Engineering foundations for distribution, for user hacking, for bringing more people into the system, these are the things that allowed open source to emerge in software. Good design choices, like separation of concerns, led us to the world of open source software. Let's learn from those lessons and build the foundations first, and let the science surprise us with the way it localizes distributed and user driven innovation. &lt;br /&gt;
&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/11/distributed_science_part_2.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/6ySN0TN9shA" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/11/distributed_science_part_2.php</guid>
         <category />
         
         <pubDate>Thu, 05 Nov 2009 07:45:41 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/11/distributed_science_part_2.php</feedburner:origLink></item>
      
      <item>
         <title>Open Source Science? Or Distributed Science?</title>
          <description>&lt;p&gt;I was asked in an interview recently about "&lt;a href="http://www.google.com/search?q=open+source+science&amp;ie=utf-8&amp;oe=utf-8&amp;aq=t&amp;rls=org.mozilla:en-US:official&amp;client=firefox-a"&gt;open source science&lt;/a&gt;" and it got me thinking about the ways that, in the "open" communities of practice, we frequently over-simplify the realities of how software like &lt;a href="http://en.wikipedia.org/wiki/Linux"&gt;GNU/Linux&lt;/a&gt; actually came to be. &lt;a href="http://en.wikipedia.org/wiki/Free_and_open_source_software"&gt;Open Source&lt;/a&gt; refers to a software worldview. It's about software development, not a universal truth that can be easily exported. And it's well worth unpacking the worldview to understand it, and then to look at the realities of open source software as they map - or more frequently do not map - to science.&lt;/p&gt;

&lt;p&gt;The foundations of open source software are relatively easy to track. In the beginning, there was free software and &lt;a href="http://en.wikipedia.org/wiki/Richard_Stallman"&gt;Richard Stallman&lt;/a&gt;. RMS didn't just invent the &lt;a href="http://en.wikipedia.org/wiki/GNU_GPL_license"&gt;GPL&lt;/a&gt; as a legal, he wrote crucial foundational software for writing software, notably the &lt;a href="http://en.wikipedia.org/wiki/GNU_Compiler_Collection"&gt;GNU compiler collection&lt;/a&gt;, &lt;a href="http://en.wikipedia.org/wiki/GNU_Debugger"&gt;GNU Debugger&lt;/a&gt;, and the original &lt;a href="http://en.wikipedia.org/wiki/Emacs"&gt;Emacs&lt;/a&gt;. So from the beginning, there was not only a free legal tool, but tools for coding that were better than other systems at the time.  &lt;/p&gt;

&lt;p&gt;Simultaneously, we can see that the emergence of &lt;a href="http://en.wikipedia.org/wiki/Microcomputer"&gt;microcomputers&lt;/a&gt; and ubiquitous access to the internet expanded the number (and interconnectivity) of potential programmers. Suddenly there were tens of thousands of programmers with computers at home and at work. The explosion of the Web saw the creation of infrastructure like &lt;a href="http://sourceforge.net/"&gt;code repositories&lt;/a&gt;, &lt;a href="http://en.wikipedia.org/wiki/List_of_revision_control_software"&gt;version control systems&lt;/a&gt;, and coding communities. Thanks to &lt;a href="http://en.wikipedia.org/wiki/Object-oriented_programming"&gt;object-orientation&lt;/a&gt;, software was also very amenable to being broken into defined, modular chunks and tasks. One coder could work on a kernel function, another on a user interface function, a third on an application, and they could be reasonably sure that as long as they all followed the standards, their work would snap together into the growing distribution. The phrase "open source" can sort of be a shorthand for this kind of innovation, which we also see in wikipedia and other community built projects. &lt;/p&gt;

&lt;p&gt;Open source, if we view it through a different lens, is really more about a &lt;a href="http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1021034"&gt;distributed methodology for software development.&lt;/a&gt; The burden of creation is widely distributed across a massive community with more-or-less equal access to tools and systems. In this context, the role of the legal tool is more akin to an enzyme. It was an essential piece of a puzzle, but it was not the only piece. In fact, without the rest of the infrastructure (connectivity, tools, and people) the legal tool on its own would not have led us to GNU/Linux. &lt;/p&gt;

&lt;p&gt;Yet far too often the focus on "porting" open source to science focuses on the legal aspects rather than performing an analysis of the infrastructure for science. Science is actually not very similar to modern software at this point. In science, especially life science, many of these factors don't exist. There isn't democratic access to tools. You tend to need a lab, which means you tend to need to work at a place big enough to afford a lab, which tends to mean you need an advanced degree, which means &lt;em&gt;there is no crowd&lt;/em&gt; - thus the fundamentals for distributed science development aren't there. And when we try to force open source on a knowledge space that is fundamentally poorly structured for distributed development, we'll not only be frustrated by our failures to replicate the GNU/Linux and Wikipedia successes, we'll risk discrediting the idea of distribution itself.&lt;/p&gt;

&lt;p&gt;Another problem: the open source approach, which is based on the open licensing of a powerful, moderately internationally harmonious property right, doesn't really apply very well to science, in which the IP situation is far more often patents v trade secret instead of copyright v copyleft. Copyrights are free to acquire, and thus easy to license at no cost as well. No one's losing an investment they made of $50,000 or more to acquire their copyright when they license code under copyleft. Patents are not so amenable to legal &lt;a href="http://en.wikipedia.org/wiki/Aikido"&gt;aikido&lt;/a&gt;. And they can kill a great idea in the cradle by tying up all the rights in a tangle of patent thickets and expensive licenses.&lt;/p&gt;

&lt;p&gt;A third problem is that science is a long, long, long, long, long way from being a modular knowledge construction discipline. Whereas writing code forces the programmer to compile the code, and the standard distribution forces a certain amount of interoperability, scientists typically write up their knowledge as narrative text. It's written for human brains, not silicon compilers. Scientists are taught to think in a reductionist fashion, asking smaller and smaller questions to prove or disprove specific hypotheses. This system almost guarantees that the tasks fail to achieve modularity like software, and also binds scientists through tradition into a culture of writing their knowledge in a word processor rather than a compiler. Until we can achieve something &lt;a href="http://neurocommons.org"&gt;akin to object-orientation in scientific discourse&lt;/a&gt;, we're unlike to see the distributed innovation erupt as it does in culture and code.&lt;/p&gt;

&lt;p&gt;A fourth problem is that science has the additional problem of collective action congestion created by the significant institutional participation impact of research institutions, tech transfer offices, venture capital, startups, and so forth. Software isn't subject to these constraints, at least, not most software. But science is like writing code in the 1950s - if you didn't work at a research institution then, you probably couldn't write code, and if you did, you were stuck with punch cards. Science is in the punch cards stage, and punch cards aren't so easy to turn into GNU/Linux.&lt;/p&gt;

&lt;p&gt;None of this is meant to discourage open approaches. We need to try. The problems we face, from neglected diseases to climate change to earthquake analysis to sustainability, are so complex that they'll probably overwhelm any approach that is not inherently distributed. Distributed systems scale much better than non-distributed, closed systems. But we should always understand the foundations, and closely examine our work to see if we need to work on building those foundations.&lt;/p&gt;

&lt;p&gt;In the sciences, the first foundation is &lt;a href="http://www.earlham.edu/~peters/fos/overview.htm"&gt;access to the narrative texts&lt;/a&gt; that form the canon of the sciences. Tens of thousands of papers are published a year. They need object-orientation - semantics - so that we can begin to treat that information as a platform, not a consumable product. Licensing is a part of this, but so is technology and scientific culture. &lt;a href="http://www.obofoundry.org/"&gt;Better ontologies&lt;/a&gt;, buy-in to technical standards, publisher participation in integration and federation, and more will be foundational to the establishment of content-as-platform. As the &lt;a href="http://research.microsoft.com/en-us/collaboration/fourthparadigm/default.aspx"&gt;data deluge&lt;/a&gt; intensifies, this foundation becomes more and more important, as the literature provides the context for the data. Moving to a &lt;a href="http://linkeddata.org/"&gt;linked web&lt;/a&gt; or &lt;a href="http://www.w3.org/2001/sw/"&gt;semantic web&lt;/a&gt; without a powerful knowledge platform at the base is building a castle made of sand - close to the water line.&lt;/p&gt;

&lt;p&gt;Another foundation is access to tools and the creation of fundamental open tools. We need the biological equivalent of the C compiler, of Emacs. Stem cells, mice, vectors, plasmids, and more will need to available outside the old boy's club that dominates modern life sciences. We need access to supercomputers that can run massive simulations for earth sciences and climate sciences. These tools need to be democratized to bring the beginning of distributed knowledge creation into labs, with the efficiencies we know from eBay and Amazon (of course, these tools should perhaps be restricted to authenticated research scientists, so that we don't get garage biologists accidentally creating a super-virus).&lt;/p&gt;

&lt;p&gt;The legal aspects weave through these foundations. The license has power to create freedoms but the improper application of a license approach carries significant risks. The "open source" meme can often feel a little religious about licenses, but it's good to remember that the GPL was invented not in the desire to write a license, but in a desire to return programming to a free state. With data and tools, we have the chance to avoid the intellectual property trap completely - if we have the nerve for it. &lt;/p&gt;

&lt;p&gt;There is some distributed innovation happening in new fields of science, like &lt;a href="http://diybio.org/"&gt;DIY biology&lt;/a&gt;, and in non science communities, like &lt;a href="http://patientslikeme.com"&gt;patients sharing treatments and outcomes with each other.&lt;/a&gt; A quick examination of the foundations reveals they are ripe for distribution: DIY biology can build on &lt;a href="http://openwetware.org/"&gt;open wetware&lt;/a&gt;, the &lt;a href="http://parts.mit.edu"&gt;registry of standard biological parts&lt;/a&gt;, and the availability of equipment and tools. Patients can connect using Web 2.0 and talk to each other without intermediaries. But this doesn't scale across into traditional science.&lt;/p&gt;

&lt;p&gt;I propose that the point of this isn't to replicate "open source" as we know it in software. The point is to create the essential foundations for distributed science so that it can emerge in a form that is locally relevant and globally impactful. We can do this. But we have to be relentless in questioning our assumptions and in discovering the interventions necessary to make this happen. We don't want to wake up in ten years and realize we missed an opportunity by focusing on the software model instead of designing an open system out of which open science might emerge on its own.&lt;br /&gt;
&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/10/open_source_science_or_distrib.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/tLekVMSBBOo" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/10/open_source_science_or_distrib.php</guid>
         <category />
         
         <pubDate>Fri, 30 Oct 2009 08:58:52 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/10/open_source_science_or_distrib.php</feedburner:origLink></item>
      
      <item>
         <title>Story Time</title>
          <description>&lt;p&gt;This post was prompted by the combination of three events: a visit with the founder of &lt;a href="http://pubget.com/"&gt;PubGet&lt;/a&gt;, an invitation to keynote at a conference on publishing, and an &lt;a href="http://scienceblogs.com/clock/2009/09/scienceonline09_-_interview_wi_12.php"&gt;interview with Bora&lt;/a&gt; about the&lt;a href="http://www.scienceonline09.com/index.php/wiki/"&gt; Science Online 2009&lt;/a&gt; conference last January in RTP.&lt;/p&gt;

&lt;p&gt;The past year has seen an explosion of talk about the future of the scientific article. It's wonderful to see, even if the results are either &lt;a href="http://www.ploscompbiol.org/doi/pcbi.1000361"&gt;depressingly complicated to achieve&lt;/a&gt; or &lt;a href="http://beta.cell.com/"&gt;depressingly incremental innovation&lt;/a&gt;. Both of those results are better than when I got into this - I remember at a &lt;a href="http://www.lub.lu.se/ncsc2006/"&gt;conference in Sweden in 2006&lt;/a&gt; hearing a grand high priest of the publishing industry argue that they'd gotten this whole digital publishing thing sorted right out...that attitude was the first thing that needed to change. Glad it has.&lt;/p&gt;

&lt;p&gt;I've been hammering for years now on the need to enrich articles with semantics. My talk at that conference in Sweden was probably the first good one I gave on the topic, and it's been an &lt;a href="http://en.wikipedia.org/wiki/Leitmotif"&gt;leitmotif&lt;/a&gt; for me going back to the mid-1990's when I was studying epistemology and getting my first real exposure to networked computers. For years I was convinced it was right around the corner. &lt;/p&gt;

&lt;p&gt;That semantic publishing future now feels closer than it ever has. But I'm actually less convinced it's around the corner than in years past, and the reasons for that are human, not technical.&lt;/p&gt;

&lt;p&gt;To be clear: in the following, I'm going to be talking about narratives and text, not about databases. The semantic future for databases and data &lt;a href="http://linkeddata.org/"&gt;is&lt;/a&gt; &lt;a href="http://www.w3.org/2001/sw/"&gt;already&lt;/a&gt; &lt;a href="http://neurocommons.org"&gt;here&lt;/a&gt;, but to paraphrase William Gibson, it's just unevenly distributed. Those of the argument that the Semantic Web isn't going to work have already lost the argument. You just don't see it, because it's an infrastructure upgrade to the back-end of the Web to make it work for data. &lt;/p&gt;

&lt;p&gt;But the impact of formal semantics on text, which is what humans interface with, has been negligible. It's had nowhere near the impact of tagging and folksonomy. That's driven me, and many others who like formal semantics, crazy.&lt;/p&gt;

&lt;p&gt;The benefits to a formal semantic approach to text are so obvious: we can start to treat knowledge as a graph, and we can even maybe start to get some network externality benefits to that knowledge. Make it more valuable via the network...one fact is like one fax machine, but many facts build a hypothesis, etc. etc. etc.&lt;/p&gt;

&lt;p&gt;Beautiful dream. Not going to happen anytime soon. &lt;/p&gt;

&lt;p&gt;The problem is that people are the writers. Humans. Not machines. Machines luuuuuv semantics. Otherwise they can't tell the difference between a picture and a pitcher (or between a pitcher of water and a baseball pitcher). This is why one should never send one's mother to buy jewelry via Google without the safe browsing mode enabled.&lt;/p&gt;

&lt;p&gt;And people don't like formal semantics. I majored in formal semantics, and it's a topic that still gives me headaches. &lt;/p&gt;

&lt;p&gt;People like stories. &lt;/p&gt;

&lt;p&gt;Scientists are people. &lt;/p&gt;

&lt;p&gt;Scientists like stories. &lt;/p&gt;

&lt;p&gt;A paper is a story. It tells, in its own way, the story of years of work. Of building expertise. Of designing falsifiable hypotheses. Of the results found in the lab. Of the search to balance those results against the canon and dogma. Of the potential ramification of the results. &lt;/p&gt;

&lt;p&gt;It's a story of science. And the telling of it is an important part of being a human who does science. &lt;/p&gt;

&lt;p&gt;A recent article in PLoS Genetics states that  "&lt;a href="http://www.plosgenetics.org/article/info:doi%2F10.1371%2Fjournal.pgen.1000622;jsessionid=BA728C02B80685DF22C9B77B025579AC"&gt;Fission Yeast Tel1ATM and Rad3ATR Promote Telomere Protection and Telomerase Recruitment&lt;/a&gt;" - now, those are the key "facts" asserted. They could be written into machine-readable format. I will spare you what that would look like. Suffice to say it's eye bleedingly ugly, and requires lots of agreement about unique identifiers. It's doable. It's being done for the databases and that will eventually make it possible for the literature. It's just not fun. And it ignores the story. &lt;/p&gt;

&lt;p&gt;It reduces the research tale to a few assertions, nested into a massive graph of stuff other people asserted. While this is great for machines, it is lousy for people. &lt;/p&gt;

&lt;p&gt;This is all leading up to an idea I'm working on for the talk later this month. Publishers need to be in the business of providing the service that translates the stories for the machines to understand. The Web makes it trivial to publish stories in human readable form. All the beautiful layout services and print services that used to be worth paying for...aren't. Peer review isn't free, but it's nowhere near as expensive as it's made out to be - and it's going to get transformed by the Web, too. The Web makes peer review massively more powerful as it makes it massively more democratic. The Web kills a lot of things that used to drive value in content, especially controlled content.&lt;/p&gt;

&lt;p&gt;After all, I can't remember the last time I used a Zagat's guide. Not when I have &lt;a href="http://chow.com"&gt;Chowhound&lt;/a&gt;. It's going to come to science. Don't know exactly how, but it's coming.&lt;/p&gt;

&lt;p&gt;But this only covers one piece of science - the telling of the story. There's another key, which is the ability to use the information to write a new tale. The ability to take this massive corpus of story and turn it into something that can be modeled, that can be used by humans and machines together to draft new stories...that ability is going to require the emergence of publishers who understand their role in the new content economy. It's not as printers who use bits rather than ink. It's as translators between the human stories and the machines who have to take those stories, integrate them into a web of linked data, and make it possible for humans to ask questions, dream dreams, and tell new stories.&lt;/p&gt;

&lt;p&gt;The semantic article isn't going to come from individual scientists rebelling and marking up their own text. It's going to be a publisher value-added service - "let us make your article integrated, and comprehensible, so that you maximize your citation count and potential collaboration." &lt;/p&gt;

&lt;p&gt;Sounds good, doesn't it?&lt;/p&gt;

&lt;p&gt;Focusing on the control of copies of the article, of the story, isn't just a losing strategy because of the open access movement, although it is that as well. It's the wrong concept entirely. Translation is a service for which authors would gladly pay. For which searchers would gladly pay. And it's a market that is going to get more valuable as a result of open systems, not less valuable, as the cost of controlled scientific published content drops thanks to &lt;a href="http://www.earlham.edu/~peters/fos/overview.htm"&gt;green and gold open access&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;Think about Clayton Christensen's &lt;a href="http://interestingprinciples.blogspot.com/2009/03/law-of-conservation-of-attractive.html"&gt;law of conservation of attractive profits&lt;/a&gt;: "When attractive profits disappear at one stage in the value chain because a product becomes commoditized, the opportunity to earn attractive profits with proprietary products usually emerges at an adjacent stage."&lt;/p&gt;

&lt;p&gt;Publishers are trying to fight the commoditization of the story. They shouldn't. The vast majority of the stories are bought and paid for by the public one way or the other. Publishers should be looking at the place where they can compete on proprietary services, and taking over those markets before their competitors - or startups - beat them to it. There is enormous opportunity in the emerging open access world to make money without needing to vigilantly police the movement of content. &lt;/p&gt;

&lt;p&gt;Help the scientists tell their stories in a way that lets those stories integrate into the digital web. Don't just gussy up a paper version of a story with hyperlinks. Don't focus on controlling the movement of stories. They're sand in your hands once they're on the network. Embrace that fact. Find the value in the next layer, the service layer. &lt;/p&gt;

&lt;p&gt;Be a guide. Be a search engine. Be a &lt;em&gt;translator&lt;/em&gt;. &lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/09/this_post_was_prompted_by.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/IB7raOLs3LM" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/09/this_post_was_prompted_by.php</guid>
         <category />
         
         <pubDate>Wed, 02 Sep 2009 12:56:54 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/09/this_post_was_prompted_by.php</feedburner:origLink></item>
      
      <item>
         <title>Ignore this post</title>
          <description>&lt;p&gt;Seriously. Just getting around to technorati claiming. Move along, nothing to see here. Watch for a lengthy post on scientific publishing later tonight or tomorrow.&lt;/p&gt;

&lt;p&gt;59tbcg4wsi&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/09/ignore_this_post.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/T83wZKcuxdY" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/09/ignore_this_post.php</guid>
         <category />
         
         <pubDate>Tue, 01 Sep 2009 19:59:30 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/09/ignore_this_post.php</feedburner:origLink></item>
      
      <item>
         <title>Open Data: It's About Interoperability, Not Property</title>
          <description>&lt;p&gt;I wrote this up on the request of a colleague who heard my talk recently on open data. I'm posting it here for comment and adding some hyperlinks...&lt;/p&gt;

&lt;p&gt;Moving from a Web of documents to a &lt;a href="http://www.w3.org/2001/sw/"&gt;Web of data&lt;/a&gt; (or of &lt;a href="http://linkeddata.org/"&gt;Linked Open Data&lt;/a&gt;) is an oft-cited goal in the sciences. The Web of data would allow us to link together disparate information from unrelated disciplines, run powerful queries, and get precise answers to complex, data-driven questions. It's an undoubtedly desirable extension of the way that the existing networks increase the value of documents and computers through connectivity - &lt;a href="http://en.wikipedia.org/wiki/Metcalfe%27s_law"&gt;Metcalfe's Law&lt;/a&gt; applied to more complex information and systems.&lt;/p&gt;

&lt;p&gt;However, making the Web of data turns out to be a deeply complex endeavor.  Data - here, a catchall word covering databases and datasets and generally meaning here information that is gathered in the sciences as a result of either experimental work or environmental observation - require a much more robust and complete set of standards to achieve the same "web" capabilities we take for granted in commerce and culture. &lt;/p&gt;

&lt;p&gt;Unlike documents, the ultimate intended reader of most data is a machine. Some classic examples include search engines, analytic software, database back ends, and more. There is simply too much data in production to place people on the front lines of analysis. When data scales easily into the petabytes, we just can't keep up using the existing systems. &lt;/p&gt;

&lt;p&gt;This machine-readability requirement is very different from the Web of documents, which was designed to standardize the way information is shown to people. Machine readability means we have to think, early and often, about the level of interoperability in any given chunk of data. "How "connectable" is it to other data?" should be the first question we ask of new data, because the level of effort required to make data connectable post-hoc is significant - frequently unbearable.&lt;/p&gt;

&lt;p&gt;The connectability quotient creates significant pressures to build interoperability deep into the Web of data. It implies a level of rigor in the design of data that understands the intended use of that data is in a network context.  Thus, we need to turn ourselves to the concept of interoperability and examine what it means in a data context.&lt;/p&gt;

&lt;p&gt;There are three interlocking dimensions to interoperability in data: legal, technical, and semantic. By legal, we mean the contractual and intellectual property rights associated with the data; by technical, the standard systems (especially the computer languages) in which the data is published; and by semantic,  the actual meaning of the data itself - what it describes, and how it relates to the broader world. &lt;/p&gt;

&lt;p&gt;Each of these dimensions is complex on its own. Taken together, the three represent unsolvable complexity. The semantic layer alone requires an almost miraculous level of agreement on "what things mean," and anyone who has witnessed argument among scientists, be they economists of physicists, knows that even apparently simple topics turn contentious over matters as basic as definitions. Consensus on the technical layer is somewhat easier - the existence of the Web and the Semantic Web "stack" of standard technologies has begun to take a leadership position in data networking - but still difficult, long, and open to argument. One of the only opportunities we have is in the legal layer, where we can look to a broad set of successes in legal interoperability through the use of a simple, flat standard: the public domain.&lt;/p&gt;

&lt;p&gt;The public domain is a very simple concept - no rights are reserved to owners, and all rights are granted to users. The public domain exists as a counterweight to copyright in the creative space, but in some countries - especially the United States - as a first option for data that is not considered "creative." &lt;/p&gt;

&lt;p&gt;The public domain option currently underpins a wide variety of linked data that is already well on its way to achieving Web scale. From the &lt;a href="http://www.ivoa.net/"&gt;International Virtual Observatory&lt;/a&gt;, whose members build an international data net on norms of "acknowledgment" rather than contracts of "attribution", to the world of genomics, where &lt;a href="http://www.ncbi.nlm.nih.gov/Genbank/"&gt;entire genomes and related data are harmonized nightly across multiple countries&lt;/a&gt;, the public domain creates complete interoperability at the legal layer of the data network, and serves as a foundation for the next layer of technical interoperability. &lt;/p&gt;

&lt;p&gt;Interestingly we have yet to observe similar network effects emerging in cases where the underlying data is treated in a more conservative "intellectual property" context by using copyright licenses or database licenses inspired by copyright. Indeed, in the case of the international consortium mapping human genomic variation, the implementation of a&lt;a href="http://www.worldlii.org/int/other/PubRL/2003/4.html"&gt; "click through" license&lt;/a&gt; was found in practice to &lt;a href="http://www.sanger.ac.uk/Info/Press/2004/041213.shtml"&gt;impede integration of that mapped variation with other public domain data&lt;/a&gt;, limiting the value of the map. The license was removed, the &lt;a href="http://www.hapmap.org/guidelines_hapmap_data.html.en"&gt;public domain option instated&lt;/a&gt;, and the database was immediately technically integrated with the rest of the international web of gene data. &lt;/p&gt;

&lt;p&gt;The legal element is of course just the beginning. The entities inside the databases themselves must be &lt;a href="http://neurocommons.org/page/URIs"&gt;named and linked&lt;/a&gt;, in a standard way. Consensus on a dizzying array of technical standards must be achieved through working groups and hard won agreement. Semantic agreement - or disagreement - must be enabled where possible, and managed through savvy technology where not possible. But if the entire system must begin with a complex set of legal terms and conditions, and be subject to the kinds of injunctions and property claims so familiar from the creative world, it is inherently unstable and unlikely to interoperate. &lt;/p&gt;

&lt;p&gt;We have seen the public domain option work, again and again, across the scientific disciplines. Implementing the public domain as the interoperability standard for the legal dimension of the web of data holds the greatest promise for scalability and long-term achievement of the network effect for data, as it permits the widest range of experimentation and development at the technical and semantic layers.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/08/open_data_its_about_interopera.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/NdAPeC0RrrY" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/08/open_data_its_about_interopera.php</guid>
         <category />
         
         <pubDate>Thu, 20 Aug 2009 11:43:31 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/08/open_data_its_about_interopera.php</feedburner:origLink></item>
      
      <item>
         <title>May All Your Standards Be Simple and Evolvable</title>
          <description>&lt;p&gt;I was in a roundtable yesterday talking about Health IT with a bunch of very smart people in the bay area. It was sort of a briefing of ourselves and others about the real issues underpinning what it would take to generate real disruptive innovation in health technology and health costs. The vast majority of the conversation centered on payment reform, which is outside my ambit. &lt;/p&gt;

&lt;p&gt;But we did spend some time talking about health data standards, and the problem of getting standards that are so geared to the existing market-dominant companies that they actually froze out new market entrants. My contribution in all this was pretty small, and to me seemed obvious. The standard that works best tends to be the least powerful solution to the problem, especially if it's an openly released solution. This can be counterintuitive - why wouldn't we want the most powerful one? - but it's been proven again and again. &lt;/p&gt;

&lt;p&gt;In technology, standards propagate like kudzu. Most of them go nowhere, representing an enormous sunk cost of time and money. And that's because most of them are way too complex. The more powerful they are, the more brittle they are, the more expensive they are to implement, and the more they restrict the re-use of the system.&lt;/p&gt;

&lt;p&gt;Tim Berners-Lee calls this the &lt;a href="http://www.w3.org/DesignIssues/Principles.html"&gt;Rule of Least Power&lt;/a&gt;, and it's one of the most important lessons I learned working at the W3C. There's a simple reason for this - the more basic the markup of the content, the easier it is to write applications that process the content. &lt;/p&gt;

&lt;p&gt;Thus TCP/IP, created simply to move bits between computers, begat a variety of new protocols like &lt;a href="http://en.wikipedia.org/wiki/File_Transfer_Protocol"&gt;FTP&lt;/a&gt;, &lt;a href="http://en.wikipedia.org/wiki/Gopher_(protocol)"&gt;Gopher&lt;/a&gt;, &lt;a href="http://en.wikipedia.org/wiki/Finger_protocol"&gt;Finger&lt;/a&gt;, many other protocols that layered atop the basic bits standard. Complexity from simplicity. Attempting to embed file transfer into the bits protocol would have made this whole process a lot harder. &lt;/p&gt;

&lt;p&gt;And of course HTML/HTTP begat the entire Web, all the way to YouTube and Amazon and everything else. Writing video codes into HMTL wouldn't have worked nearly as well as writing a standard that was simple enough to be extended by smart users coming along ten years later.&lt;/p&gt;

&lt;p&gt;To the rule of least power we can add the rule of openness - the standards process should be as open as is feasible, and the standards themselves must be open. Users have to be able to read a standard, and to have the freedom to implement the standard, to be able to innovate atop it with new systems. &lt;/p&gt;

&lt;p&gt;There's a lesson here. Gathering the relevant powers that be to figure out a standard is an important task. The &lt;a href="http://www.w3.org/"&gt;W3C&lt;/a&gt;, the &lt;a href="http://www.ietf.org/"&gt;IETF&lt;/a&gt;, the &lt;a href="http://www.omg.org/"&gt;OMG&lt;/a&gt; (that's Object Management Group, not the internet acrony, for you younguns), and what &lt;a href="http://www.google.com/search?client=safari&amp;rls=en-us&amp;q=data+standards&amp;ie=UTF-8&amp;oe=UTF-8"&gt;feels like every different data discipline&lt;/a&gt; on earth does standards this way.&lt;/p&gt;

&lt;p&gt;But there's a lot of fingers on the scale for most of this work. That's because data standards tend to get created by well-meaning, overworked, and underpaid people who are making a real sacrifice to work on the standards. And those people are going to depend on a lot of in-kind work from the interested parties, who are always going to try to bend the standards to their will. &lt;/p&gt;

&lt;p&gt;That can go multiple ways. The paranoid conclusion is that the for-profits involved will try to use the standard to increase stock prices, which is why smart standards efforts include &lt;a href="http://www.w3.org/Consortium/Patent-Policy-20040205/"&gt;patent policies&lt;/a&gt; to prevent enclosure. But there's a bigger problem out there, which is much less visible but much more of a force in the creation of standards that don't get used, or that don't do what we want them to do.&lt;/p&gt;

&lt;p&gt;It's what I call the problem of standards completeness. Experts in the field, interested parties, impassioned volunteers - these people by their nature tend to want to make the standard they build as complete as possible. They want to cover the most ground with the standard. They understand the space so well that they want to build standards that address vast swaths of work. &lt;/p&gt;

&lt;p&gt;But that violates the Rule of Least Power. And as we move towards a web of data, even a &lt;a href="http://www.nhinwatch.com/"&gt;web of patient data&lt;/a&gt;, we'll do well to make our standards by solving real problems with the simplest possible solutions, then releasing those solutions for others to build on. &lt;/p&gt;

&lt;p&gt;The impact of the simple evolvable standard in short term is probably less than a more complete, perfect standard. Certainly TCP/IP didn't scare the systems integrators at its inception. But it's the power of the crowd that can build on the open standard that breaks open the market. Thanks to simple standards, &lt;a href="http://en.wikipedia.org/wiki/Google#History"&gt;two talented programmers can start a company in a garage that changes the world&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;If we're going to bring that level of innovation potential to health IT, we need to keep the lessons of the simple standard in mind. Because right now, if you're a bright young entrepreneur, you don't get into health IT. And the lack of not just standards, but the right kinds of standards, is the first barrier we have to knock down to change that reality.&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/08/things_you_dont_want_to_watch.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/V9gLOrDTW4g" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/08/things_you_dont_want_to_watch.php</guid>
         <category />
         
         <pubDate>Wed, 05 Aug 2009 12:05:36 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/08/things_you_dont_want_to_watch.php</feedburner:origLink></item>
      
      <item>
         <title>Integrate. Annotate. Federate. </title>
          <description>&lt;p&gt;Following on to &lt;a href="http://scienceblogs.com/commonknowledge/2009/07/publishing_science_on_the_web.php"&gt;yesterday's post&lt;/a&gt;, where I wrote about the four functions that traditional publishers claim as their space (registration, certification, dissemination, preservation), I want to revisit an argument I made last week at the British Library. &lt;/p&gt;

&lt;p&gt;In &lt;a href="http://www.slideshare.net/wilbanks/future-of-scientific-communication-what-is-the-genuine-article"&gt;my slides&lt;/a&gt;, I argued that the web brings us at least three additional functions: integration, annotation, and federation. I wanted to get this argument out onto the web and get some feedback...&lt;/p&gt;

&lt;p&gt;Let's start with integration. The article no longer sits on a piece of dead tree, inside a journal formatted by date and volume and page number. It exists as a digital entity, capable of dense integration into other digital entities. One way to think of this is to think of how the citation is truly weak tea compared to the hyperlink - an individual citation carries more weight than an individual hyperlink, but the hyperlink is so easy to create, and carries &lt;a href="http://en.wikipedia.org/wiki/PageRank"&gt;so much power in aggregate&lt;/a&gt;, that we get Google. Citations are the only way most articles are integrated with other articles, and that simply has to change.&lt;/p&gt;

&lt;p&gt;Articles need to be integrated with lots of other digital information. Media is an obvious one, and the Elsevier-Cell &lt;a href="http://beta.cell.com/erickson/"&gt;"article of the future"&lt;/a&gt; seems to start here with an interview with the authors. To me this is absurd, and the height of how a "big company" thinks "the users" use the web. I don't want to hear an author interview with a reporter. I assume the author is going to say his or her work is sweets and sparkles and Nobel prizes. I'd rather see an embedded high-resolution video of all protocols necessary to replicate the experiment like the ones you get from &lt;a href="http://www.jove.com/"&gt;JoVE&lt;/a&gt; (I'd like them to actually be open access too, but that's a &lt;a href="http://scienceblogs.com/commonknowledge/2009/04/jove_goes_closed_access.php"&gt;different blog post)&lt;/a&gt;. &lt;/p&gt;

&lt;p&gt;If you want to make the article of the future, start with integration and work backwards. Don't start with the article and work forward, because you'll be trapped in document mentality instead of the network mentality.&lt;/p&gt;

&lt;p&gt;We don't just want the data downloadable, we want to be able to run the same algorithms the author ran on the data, and adjust the variables myself, to see if the results are the output of statistical foul play or negligence. We want to be able to hide all the boring language that recapitulates past canon and focus on the new assertions, unless of course the author is trying to game the past canon and shade the facts. And we want to be able to effortlessly click out and get data about the assertions in the paper from other databases - when there's a gene mentioned, we should be able to one-click and run any number of core queries against the sequence, the ontological classifications, order genetic materials from biobanks and so forth.&lt;/p&gt;

&lt;p&gt;Annotation is the second new essential function. The old method of annotation is through either writing a new paper that validates, invalidates, extends, or otherwise affects the assertions made in an old paper. Or if something is really wrong, there might be a letter to the editor or a retraction. In a wiki world, this is fundamentally insane. The paper is a snapshot of years of incremental knowledge progress. We have much better technology to use than dead trees. &lt;/p&gt;

&lt;p&gt;Of course, there isn't any incentive to take the wiki that is science and actually use a wiki to create and edit it. Scientists get tenure for papers, and &lt;a href="http://en.wikipedia.org/wiki/Egoboo"&gt;egoboo&lt;/a&gt; is cold comfort. Annotation needs to be provided by publishers, and is being provided, but the next step is to create an open platform that actually tracks the kind of annotation-relationships that the web enables. Bloggers use &lt;a href="http://en.wikipedia.org/wiki/Trackback"&gt;trackback&lt;/a&gt; to create a formal hyperlink between blog posts, and the protocol can and should be extended to let us connect all sorts of things: articles, wiki pages, database entries, catalog pages for biological materials, data sets, and on and on. By making these link transactions - which exist anyway - explicit and trackable, and most importantly reportable, we'll create a currency that scientists will gladly spend. It won't be about "sharing" but instead about "publishing" more of the intermediate knowledge that currently gets left on the lab floor when the paper gets written.&lt;/p&gt;

&lt;p&gt;Federation is the last essential new function I'll deal with here (have some theories on other long term essential ones, but they're poorly formed in comparison). By federation I mean the ability to take a set of articles and federate them into a corpus with other materials. There's a lot of reasons one might want to do this: text mining, semantic indexing, integration with information that is private, and so forth. It's great to be able to read articles on the web. But if we're going to really explode the way we communicate, the ability to cache local copies (or cloud copies) in new formats for new kinds of analysis, and the right to then distribute the resulting corpus for follow-on innovation and exploration, is going to be central. &lt;/p&gt;

&lt;p&gt;Publishers are so focused on the prevention of copying that they don't see the central business opportunity here: the human-readable, copyrighted version of the article is the least federation-friendly. Charge a fee to make the article beautifully machine-readable and give away the text - because the service of improving the technical aspects of the article is clearly a value-add that shouldn't be subject to a funder mandate.&lt;/p&gt;

&lt;p&gt;Integration, Annotation, Federation. It's what the Web is all about. And if we can get to the point where publishers feel these as core responsibilities, the Open Access debate will have made a major leap. All of these create a world in which the text of the article itself is lower in economic value, and thus easily distributable, than the connectivity of that article into a larger web of information. OA is the beginning, not the end game, of making the web work for science the way it works for culture. Step two is all about the connectivity, and it's time to start arguing - loudly - for the right to start wiring the science together.&lt;br /&gt;
&lt;/p&gt; &lt;a href="http://scienceblogs.com/commonknowledge/2009/07/integrate_annotate_federate.php#commentsArea"&gt;Read the comments on this post...&lt;/a&gt;&lt;img src="http://feeds.feedburner.com/~r/scienceblogs/CommonKnowledge/~4/iiaKvTx-jic" height="1" width="1"/&gt;</description>
         <guid isPermaLink="false">http://scienceblogs.com/commonknowledge/2009/07/integrate_annotate_federate.php</guid>
         <category />
         
         <pubDate>Fri, 31 Jul 2009 11:21:22 -0500</pubDate>
      <feedburner:origLink>http://scienceblogs.com/commonknowledge/2009/07/integrate_annotate_federate.php</feedburner:origLink></item>
      
   </channel>
</rss>