I've been wrestling with a wishful thinking manuscript, and of course -- as always seem to be the case -- the data are getting in the way of a good theory. Wishful thinking is when preference leads to expectation. In other words, if I'm for Candidate X, I'm probably also likely to predict Candidate X will win the election. How often does it happen? See below. This graphic I created from the ANES cumulative data file and goes only to 2004, but I'll tell ya now the wishful thinking in 2008 was just as high. The predictive accuracy is just that, what percentage accurately predicted the election. As can be seen, in close elections the accuracy tends to be low, in runaway elections, it tends to be higher.
I may write up a version for Like the Dew or one of my other favorite places. The paper I'm working on takes a simple idea -- that those who expect to win, but don't, are likely to have lower scores in government trust or satisfaction with democracy. And, from the media angle, exposure to the news should moderate this effect.
But do the data show that? Not as well as I'd like, so I may push this idea aside.
Random blog posts about research in political communication, how people learn or don't learn from the media, why it all matters -- plus other stuff that interests me. It's my blog, after all. I can do what I want.
Showing posts with label data analysis. Show all posts
Showing posts with label data analysis. Show all posts
Monday, March 15, 2010
Friday, September 11, 2009
Social Networks -- as Predictors of Popularity
This appeals to my inner geek, the idea of using social networks as a way to predict the next big hit in pop music.
This article (warning: math) asks whether it's better to rely in the intrinsic qualities of music or on social networking to predict that next song to run up the charts. I won't get into the math but the methodology looks sound to me. The authors used last.fm as source material for user tags to create a "folksonomy of music."
The result? Mining the music social networking data does a pretty damn good job of identifying popular songs. Cool.
Why am I so fascinated?
Think of the ways we might mine Twitter or Facebook or whatever tomorrow's social networking flavor of the year happens to be to predict the success of a policy, like health care, or even of a candidate's run for office. If we take social networks to roughly sum up the group mind, then tapping into that should be an interesting snapshot -- not unlike that of a traditional public opinion poll. I see this as another way of understanding what people think and know about the important, and not-so-important, issues of the day.
Yeah, the math and algorithms are not for the casual user. Gotta work out the kinks, but I suspect someone will come up with off-the-shelf and on-the-net ways for paid users to do this kind of analysis.
This article (warning: math) asks whether it's better to rely in the intrinsic qualities of music or on social networking to predict that next song to run up the charts. I won't get into the math but the methodology looks sound to me. The authors used last.fm as source material for user tags to create a "folksonomy of music."
The result? Mining the music social networking data does a pretty damn good job of identifying popular songs. Cool.
Why am I so fascinated?
Think of the ways we might mine Twitter or Facebook or whatever tomorrow's social networking flavor of the year happens to be to predict the success of a policy, like health care, or even of a candidate's run for office. If we take social networks to roughly sum up the group mind, then tapping into that should be an interesting snapshot -- not unlike that of a traditional public opinion poll. I see this as another way of understanding what people think and know about the important, and not-so-important, issues of the day.
Yeah, the math and algorithms are not for the casual user. Gotta work out the kinks, but I suspect someone will come up with off-the-shelf and on-the-net ways for paid users to do this kind of analysis.
Labels:
data analysis,
facebook,
music,
social networks,
twitter
Thursday, September 10, 2009
Song Lyrics of the Day
A Don Henley song, The Garden of Allah, includes this bit of lyrics:
Kinda sums up my scholarly approach ...
Because there are no facts, there is no truth,
just data to be manipulated
I can get any result you like
What's it worth to you?
Kinda sums up my scholarly approach ...
Thursday, August 6, 2009
Cool Stuff: MemeTracker
Check out MemeTracker, a neat visualization of news coverage. Scroll down the first page to see some of the other applications with these data.
This is all part of the growing field of statistically analyzing the boatloads of data out there to help make sense of what we do, what we think, and how we respond on the Net. And it's just fun to play with. I like the "top phrases" graph. Run your mouse over the graphs and the phrase pops out. Cool.
What's this all mean? Lots of stuff, but it'll be an interesting portrait of what the media are talking about, what bloggers are talking about, what news people are talking about, another piece of the puzzle along with opinion polls and all the rest. What this fails to do, of course, is get at the meat of the matter. It's a picture, a snapshot, but that's about it. We don't so much learn from this as we see what we've been talking about. Still, a lot of fun to play with, and the first of many steps that hopefully will take us from what we're talking about to what we know.
This is all part of the growing field of statistically analyzing the boatloads of data out there to help make sense of what we do, what we think, and how we respond on the Net. And it's just fun to play with. I like the "top phrases" graph. Run your mouse over the graphs and the phrase pops out. Cool.
What's this all mean? Lots of stuff, but it'll be an interesting portrait of what the media are talking about, what bloggers are talking about, what news people are talking about, another piece of the puzzle along with opinion polls and all the rest. What this fails to do, of course, is get at the meat of the matter. It's a picture, a snapshot, but that's about it. We don't so much learn from this as we see what we've been talking about. Still, a lot of fun to play with, and the first of many steps that hopefully will take us from what we're talking about to what we know.
Wednesday, May 27, 2009
Rejection
Nothing like a journal rejection to start your day, this on a manuscript so much better than a previous one the same journal had accepted with only minor revisions and quickly published. Gotta love the crapshoot that is peer review. Time to rework it then shop the thing around in hopes of finding it a good home. Already have a target in mind.
In other off-topic news, working on my Public Opinion seminar syllabus. Scouring the net for good stuff and trying to plot out my day-to-day discussion topics. It's a summer seminar that meets too many days in a week for too many hours in a day.
In yet more off-topic news, in about an hour it's Manchester United vs. Barcelona. Go Barca!
And in a final bit of off-topic, off-blog news, my data analysis on a new manuscript absolutely sucks. Not sure what to do -- toss the zillion hours of work I've put into it, or try to salvage something? It's on a topic I love (wishful thinking and predictive accuracy), but the data simply are not cooperating. Not sure what to do.
In other off-topic news, working on my Public Opinion seminar syllabus. Scouring the net for good stuff and trying to plot out my day-to-day discussion topics. It's a summer seminar that meets too many days in a week for too many hours in a day.
In yet more off-topic news, in about an hour it's Manchester United vs. Barcelona. Go Barca!
And in a final bit of off-topic, off-blog news, my data analysis on a new manuscript absolutely sucks. Not sure what to do -- toss the zillion hours of work I've put into it, or try to salvage something? It's on a topic I love (wishful thinking and predictive accuracy), but the data simply are not cooperating. Not sure what to do.
Labels:
data analysis,
journal rejection,
peer review,
public opinion
Tuesday, February 17, 2009
When Knowledge Items Go Bad
I'm working with a large data set now that includes several measures of political knowledge. One set of questions gets at what people know of the two candidates (their religion, state of residence, demographic stuff). The Cronbach's Alpha for this index is not great, about .55. An index of standard political knowledge items such as how long is a U.S. senator's term, how many times can a president be elected, standard textbook civics class stuff, that is only .60 or so. Not good.
But questions about the candidates stands on issues, that alpha is awful, somewhere between .40 and .50 depending on which items I use. These are way too low and, to be honest, unexplainable.
Oh it gets better. These indices, created quick and dirty to peek under the hood of my data, don't line up at all with my key independent variable. And dammit they should. This is a great theory with the data getting in the way (I'll discuss more fully at another time, closer to subbing this thing to a journal). Basically, there should some relationship here.
I've blown two long afternoons on massive data recoding and analysis, triple checking all my recodes.
I shoulda been a plumber. Named Joe.
But questions about the candidates stands on issues, that alpha is awful, somewhere between .40 and .50 depending on which items I use. These are way too low and, to be honest, unexplainable.
Oh it gets better. These indices, created quick and dirty to peek under the hood of my data, don't line up at all with my key independent variable. And dammit they should. This is a great theory with the data getting in the way (I'll discuss more fully at another time, closer to subbing this thing to a journal). Basically, there should some relationship here.
I've blown two long afternoons on massive data recoding and analysis, triple checking all my recodes.
I shoulda been a plumber. Named Joe.
Sunday, February 1, 2009
The New Data are Here!!!
The ANES released its 2008-2009 panel study this weekend.
Data hounds, get to sniffing!
This is a 21-month set of surveys that I only just downloaded to my office computer a few moments ago. The data supposedly include political and non-political items. I've only glanced at the codebook. A raw frequency result is here, but it's not easy reading for the uninitiated.
Keep in mind this is an advanced release. Cleaner versions usually follow. And the ANES standard 2008 pre- and post-election surveys will be released near the end of the month.
Drivers, start your SPSS . . .
Data hounds, get to sniffing!
This is a 21-month set of surveys that I only just downloaded to my office computer a few moments ago. The data supposedly include political and non-political items. I've only glanced at the codebook. A raw frequency result is here, but it's not easy reading for the uninitiated.
Keep in mind this is an advanced release. Cleaner versions usually follow. And the ANES standard 2008 pre- and post-election surveys will be released near the end of the month.
Drivers, start your SPSS . . .
Thursday, November 20, 2008
ANES and Knowledge
As we gear up for post-election analyses, a reminder that one of the dominant data sets used by many scholars has a problem in how political knowledge was measured.
This report outlines some of the issues, how they've been dealt with, or how to correct for them in analysis. I blogged about this quite some time ago but it bears repeating since, soon, a new set of ANES data will be available for those wanting to study the 2008 presidential campaign. I'm convinced the new data release will be clean and address all the issues in this report.
No official word on when the ANES data will be available. They released an early version of the 2004 election data on Jan. 31, 2005 (early meaning no coding of the various open-ended questions). In April a more full version appeared. This excludes various errata, corrections, and other tweaks that happen along the way.
Get those SPSS engines tuned up and ready to rumble.
This report outlines some of the issues, how they've been dealt with, or how to correct for them in analysis. I blogged about this quite some time ago but it bears repeating since, soon, a new set of ANES data will be available for those wanting to study the 2008 presidential campaign. I'm convinced the new data release will be clean and address all the issues in this report.
No official word on when the ANES data will be available. They released an early version of the 2004 election data on Jan. 31, 2005 (early meaning no coding of the various open-ended questions). In April a more full version appeared. This excludes various errata, corrections, and other tweaks that happen along the way.
Get those SPSS engines tuned up and ready to rumble.
Scholars will use lots of other data, but political scientists in particular love the ANES surveys. Media scholars? Not so much. The media variables tend to be less compelling, but that's another post for another day.
Labels:
2008 research,
ANES,
data analysis,
political knowledge
Subscribe to:
Posts (Atom)
