Showing posts with label databases. Show all posts
Showing posts with label databases. Show all posts

Saturday, October 25, 2008

Searching dirty

Librarian Genie Tyburski teaches an online class through the University of North Carolina on web research.

Sounds dull and dry, right? Uh, no.

Try, for example, using some of her suggested search terms on Google to find stuff that people don't want you to find:


  • "not for public dissemination"

  • "not for public release"

  • "official use only" (variations include FOUO and U//FOUO)

  • "company confidential"

  • "internal use only"



For Genie's class, I tried some of these search terms combined with *NC*. I found a local political candidate's profile, with his home phone number, marked "not for public release."

That's one general theme that emerges from Genie's class: Information that is supposed to be private can sometimes inadvertently leak onto the web, through careless coding, or scanning, or editing, or incorrect placement on a server.

We've debated in class the legality and ethics of finding such information, and concluded that using the tools to find such information fall into legal, ethical realms like much of the reporting labeled "investigative." The ethical questions get sticky when you weigh what to do with the found private information.

Regardless, Genie's tools should be familiar tools for reporters and other journalists. Read her article and play around with the search terms sometime.

And check the updated "Reference" sidebar here for links to other resources, including Power Googling.

Wednesday, June 11, 2008

Interactivity

Sometimes, you just have to "see" a story to get an idea of its scope. Merging video, text, photos and graphics gives such a richer story than just reading an article or watching a video.

StarTribune.com proved the power of this with its "13 Seconds in August" project looking at the I-35 collapse.

Now, the folks at DesMoinesRegister.com have shown us -- in ways made possible only by using the power of the Web -- the terrible damage to a town destroyed by a ferocious tornado. The Des Moines team has assembled graphics, photos, data and stories into an amazingly interactive package that's worthy of emulation.

Charles Apple, at VisualEditors.com, quotes the Registor's data editor, James Wilkerson, about how the project came together:

For base data about the properties, I scraped the county assessor’s web site using a perl script and put the results in a spreadsheet. There were about a thousand records for Parkersburg.

One of our graphics people, Kelli Morris, walked the route of the storm, taking pictures and talking to survivors. She then used the property spreadsheet to link to “after” pictures and built a library of survivor stories from her data and stories we published. We later went through all of the properties for which Kelli had pictures and downloaded the “before” photo from the assessor’s web site.

A parcel map shapefile was not available in a timely and affordable manner. So another graphics person — Craig Johnson — built his by hand. He then built the Flash display using some dummy xml data. Putting Kelli’s spreadsheet in mySQL, I built an xml page in php, which was used to fuel the final display.

The end product is something I believe is truly unique and visually powerful. It also shows what can be accomplished by graphics folks who understand how to use data and think ahead about how to best weave it into their work.


Wonderful example of the power of merging data and multimedia.

Sunday, June 8, 2008

Connecting the dots

The lead story in The Observer today was a good old-fashioned pork-barrel project with a new twist. The Associated Press Managing Editors Association created a project to train reporters in how to use new online research tools through the Sunlight Foundation. The Observer's McClatchy's Lisa Zagaroli participated and reported on government pork projects in North Carolina.
It's a step forward for the group of newspaper leaders, especially since the Sunlight Foundation is led and funded by new media people: Craig Newmark of Craigslist, Jimmy Wales of Wikipedia, and Pierre Omidyar, founder of EBay.
We live in interesting times.
And a second connect-the-dot: The Observer's Forrest Brown spoke Saturday in Greensboro for the Society of Professional Journalists' Citizen Journalism Academy, about reporting and writing basics. I haven't talked with him about it yet, but have it on strong (Twitterfriend) authority that he was good and funny. I can't wait to hear more.
So connect the dots, and think about the nonprofit, nonpartisan Sunlight Foundation's mission. They seek to use:
“new information technology to enable citizens to learn more about what Congress and their elected representatives are doing, and thus help reduce corruption, ensure greater transparency and accountability by government, and foster public trust in the vital institutions of democracy. We are unique in that technology and the power of the Internet are at the core of every one of our efforts."

And another thing: A relationship button?

Thursday, February 14, 2008

Kansas City wins at SND


Congrats to sister paper The Kansas City Star for winning an award of excellence in SND judging in Syracuse.
Results for other papers are unclear at the moment: the SND judges are designers, not database experts, and they want some time before posting the full results database. They're great at posting pictures though.
Please note: Kansas City uses CCI to produce their paper. The winning front page uses a tried-and-true formula for breaking news: great photos, played well; a locater map, a breakout box. What looks like the planned centerpiece was squished downpage while retaining its graphic elements.
Six of the Top 10 list of winners use CCI: the Los Angeles Times; The New York Times; The Boston Globe; Hartford Courant; Chicago Tribune, and the San Jose Mercury News.
Thought for next year's SND judging: invite a database geek or two to help get the full results posted faster. It's possible, as The New York Times demonstrates with election results.
Or develop one from within. Avoid fields; jump fences.

Tuesday, January 22, 2008

Lessons and links from science bloggers

Bloggers from the 2008 N.C. Science Blogging Conference have returned from last weekend's chilly event to their homes and keyboards, sharing their presentations, photos and thoughts online. What's so cool is that anyone can continue to learn from this conference from the comfort of their own monitors, wherever they might be, and whenever they can.

In addition, the conference has a heavy dose of participation from the
smart minds at Science Blogs, supported in part by Seed magazine. The networked circle of science represents one new way of aggregating and filtering information beyond the traditional methods of big-company media sites. NYU media professor Jeff Jarvis has made much of Glam for doing the same thing (perhaps with a larger emphasis on advertising and content that attracts ads). Glam doesn't impress me; its college fashion blog can't hold a candle to the Daily Tar Heel's The Good, The Bad and The Fab.

Oh but wait. We were talking journalism and science, not fashion.

I respectfully submit that Science Blogs serves as a better model for distributing, sifting and making findable strong content than sites like Glam. Ads play a supporting role, rather than being the goal.

And conference organizers are also demonstrating a new model of sharing strong content with "reverse publishing," creating a downloadable or paperback book of the best science blog posts of 2007. You can read the background of how the idea came to be at Bora Zivkovic's A Blog Around the Clock. The "publisher" of the compilation is Lulu, and the editors are Zivkovic and Reed Cartright, with input from the readers of Zivkovic's blog.

But back to the conference. The main jumping-off point of the group is a wiki.
Below are random links gleaned from various conference bloggers. They're filtered through what I find interesting and not too far above my head. Most bear a relationship to journalism; some don't. Of course, your mileage may vary.

How to report scientific research to a general audience
From Cognitive Daily, out of Davidson.
(Alternate title: How to report anything to a general audience.)
My favorite line:
"Visuals need the same treatment as words."

Peer-reviewed research
(How to dig through all the crap to find the ponies.)
(Or how wearing a badge can change the life of your blog.)
Bloggers for Peer-Reviewed Research
Research Blogging


Citizen Science

(Who knew? We thought it was all about us, the journalists and citizen journalists.)
Purple loosestrife detectives and reporters, at the U.S. Geologic Survey.
Cornell University's Citizen Science toolkit from a citizen science conference.

Public Library of Science and one of its online peer-reviewed journals, PLOS One.
Again, who knew?

A ring of science blogs
(Fix a big cup of something and stay awhile).

Politics
Questions for the next president.

PDF organization
Organize all your pdfs and papers as if they were songs on Itunes. Unbelievably valuable for people in distance-learning classes, but only if they're smart enough to have Macs. At Papers.


The Institute for Southern Studies

I remember this group from my days as a student journalist in Georgia. It's great to see they're still producing research galore. One of their most recent reports is about the "devastating costs" North Carolina is suffering from war, and it comes after the launch of the N.C. Military Foundation, a public-private entity to lure more defense contracts to North Carolina.
This site is worth digging into, keeping in mind the organization does have an agenda. It's intriguing to think about comparing its research with that available through The Sunlight Foundation and Taxpayers for Common Sense at Earmark Watch Dot Org.

Flickr groups to identify plants
(These will change my life and possibly put my aunt, The Plant Oracle of the Mountains, out of business).
ID Please and What plant is that?

Invasive species blog
(And you thought mussels only stopped development.)
Invasive Species.

Tuesday, January 8, 2008

Two cups data, one cup journalism

Matt Waite of Politifact is a smart guy. Go read him. Part of his latest:

"Are we really building a business model, or even a component of a business model, around making public data searchable? Because guess what? Google is too. That’s right. The search giant is dealing directly with government agencies to help them make their own data searchable. Sound familiar? Think your data ghetto can compete with Google? Do you think people are going to remember your newspaper.com url over Google? Really?"

"....That said, here’s how we can get out of the data ghetto: add some journalism to it."

Like Charlotte did here.

More on Google's efforts from my UNC class research last semester:
"The search engine company has launched technology and standards to make public records more findable on the Internet and is making agreements with states to help get public information in to the hands of the public. The most recent agreement was with the state of Florida, opening records about public schools, water and waste permits, employment data and consumers' commuting patterns. Google is offering its services for free for now. ...
Google has also initiated agreements with plainlanguage.gov hosted by the Federal Aviation Administration, and the Energy Department's Office of Scientific and Technical Information and the Education Department's National Center for Education Statistics." Reference.

More:
Ensuring government is only one search away here.
Agencies work with Google here.
Dense code but clues to the future here.
Hiding in plain sight: Why important government information cannot be found through commercial search engines (again, density warning): here.

H/T to Waite's post from John Hassell, from the Facebook group "The Exploding Newsroom," now a blog.

Monday, November 19, 2007

Local calendars: Our neighbors want them

As the busy holiday and art seasons approach, and as the political year of 2008 nears, our neighbors are crying out for decent social calendar tools.
Someone, or several someones, will figure out the right tools, and then steal our ads or find another way to make money from it, if we're not there first.
I could include many links of what some papers are doing, but the reality seems to be no one has it right yet. Let's keep looking and sharing.
If we share information, we can evolve faster.

Monday, October 29, 2007

Bike versus car (And you can too!)


Apologies to Mr. Colbert on the title.

Wonderful visual presentation of bike-car wrecks from The Oregonian.

We could do it in Charlotte and Raleigh, with information at sites like this one and this one and this one.

I spent about two minutes at the last one and made the (not so pretty) chart shown here. The hard part: The crash data site says, "For a detailed review of crashes in specific locations (e.g., corridors or certain intersections within a community), it will be necessary to obtain such information at the local level."

But it's possible.

Here's another example of mashing up publicly available information to make it useful for readers, from Marc Matteo with McClatchy in California.

Utopian ideas: Foundations could give rewards or grants to creators of such projects so they can take the time to write project "cookbooks" for other papers. Or foundations could fund time for local reporters and graphic artists to develop their own projects, without eviscerating slim newspaper staffs. The ideas are spinning off Ed Wasserman's critique of Propublica.

Tuesday, June 12, 2007

Who needs a full time web developer?

Hire a company to do it for you.

I had an interview today with Jean Dubail, the AME/Online at the Cleveland Plain Dealer. They've been working with a California-based company called Caspio. Basic idea is that you get them your data through a wizard on their servers. They will then generate html code that you import onto your pagees. Readers can then search the database from your site, pulling information from their servers. Cleveland.com has used this a great deal for campaign finance information, crime, etc. Seems like a great way to get the developer in your hands when you need them most.

Couple examples built with the help of this Caspio technology:
Guitarmania http://www.cleveland.com/guitarmania/
A United Way fundraiser - like the cows or the rocking chairs Charlotte had a while back. Start from this database, pick the guitar you want to look at and it pulls the vitals from the database and then maps the location for you.

The King James Statistical Bible http://www.cleveland.com/sports/lebron_stats/
One of Dubail's favorites, this one lets you take a look at Lebron's stats in any number of ways.

Earthquake report http://www.cleveland.com/weather/earthquake_report/
A minor earthquake a couple months ago. The website asked readers to add their information to the database. It's then mapped on a Google map.

Cleveland.com's top hits are routinely sports - with the Cavs in the NBA Finals that's off the charts right now. Browns fans can't get enough. Dubail gave me today's stats: the top 5 stories were all sports. About half of Monday's unique visitors went to sports pages and accounted for about 45% of the page views.
They're developing a database of all things Browns that will encompass the 60-some year history of the team. It's taking a lot of manpower, but they're aiming to have it running in time for preseason. The hope is this database will let you search game by game for stats AND for stories from the PD. There's also one in the works for a huge high school sports database of schedules, records, results and the like.
For a town where sports traffic means big business, these two seem like fabulous ideas.