| Function | Input | Output | Best for |
| apply | Rectangular | Rectangular or vector | Applying function to rows or columns |
| lapply | Anything | List | Non-trivial operations on almost any data type |
| sapply | Anything | Simplified (if possible) or list | Same as lapply, but with simplified output |
| plyr::ddply | data.frame | data.frame | Applying function to groupings defined by variables |
Showing posts with label R. Show all posts
Showing posts with label R. Show all posts
Thursday, March 14, 2013
Apply-style commands in R
Here's a quick table of what I think are the most useful apply-style commands in R:
For alternatives to plyr, see this post on StackOverflow.
Saturday, February 2, 2013
R scripts for analyzing survey data
Another site pops up with open code for analyzing public survey data:
http://www.asdfree.com/
It will be interesting to see whether this gets used by the general public--given the growing trend of data journalism and so forth--versus academics. It is a useful resource for both.
Tuesday, July 10, 2012
This is *huge*: SAScii package
http://blog.revolutionanalytics.com/2012/07/importing-public-data-with-sas-instructions-into-r.html
Q: How do you make a hairless primate?
Answer 1: Take a hairy primate, wait a few million years and see if Darwin was right.
Answer 2: Make them work in SAS and watch them pull all their hair out.
Unfortunately many public datasets are released as ASCII files with only SAS code to read them in, name all the variables properly, etc.
Now there's a new kid on the block, the SAScii package for R, which will read in the SAS script, parse it, and deliver you an R file instead. Since R has fabulous import/export abilities (via the foreign package), this means even if you are a Stata user you can take advantage.
Friday, March 16, 2012
Dates and times in R
Nothing looks funnier than a patchy simian. That's why we sighed a great sigh of relief when we spotted this article on the lubridate package in R. It saves a great deal of hair pulling.
http://www.r-statistics.com/2012/03/do-more-with-dates-and-times-in-r-with-lubridate-1-1-0/
http://www.r-statistics.com/2012/03/do-more-with-dates-and-times-in-r-with-lubridate-1-1-0/
Sunday, February 19, 2012
Reading huge files into R
SAS is much touted for its ability to read in huge datasets, and rightly so. However, that ability comes at a cost: for smaller datasets, since files remain on the disk rather than in memory (as is the case with Stata and R), it is potentially less fast.
If you don't want to learn/buy SAS but you have some large files you need to cut down to size (rather like the gorilla in the corner), R has several packages which can help. In particular, sqldf and ff both have methods to read in large CSV files. More advice is available here: http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r
If you're a Stata person, you can often get by reading the .dta file in chunks within a loop by adding e.g. "in 1/1000" afterwards if you want to read the 1st through 1000th observation in.
If you don't want to learn/buy SAS but you have some large files you need to cut down to size (rather like the gorilla in the corner), R has several packages which can help. In particular, sqldf and ff both have methods to read in large CSV files. More advice is available here: http://stackoverflow.com/questions/1727772/quickly-reading-very-large-tables-as-dataframes-in-r
If you're a Stata person, you can often get by reading the .dta file in chunks within a loop by adding e.g. "in 1/1000" afterwards if you want to read the 1st through 1000th observation in.
Thursday, June 9, 2011
We are now a part of R-bloggers.com
R-bloggers is a site that aggregates many of the best R blogs on the internet. We're glad they've allowed our R-related posts to be aggregated there. If you mainly write in R, it's worth checking them out.
R: Speeding things up
R is many things, but it's not exactly speedy like a Patas Monkey. In fact, while it is much faster than many other solutions, R is notably slower than Stata (even inspiring talks that it should be rewritten from scratch!).
Fortunately, Radford Neal has been hard at work speeding R up, and has released some new patches to play with if you find it too slow. You can also try writing key sections in C++, or using Revolution Analytics' offerings (free for academics).
For extreme speed needs, however, R can't be beat, as it has long offered graphics-card based extreme parallelism that commercial solutions are only beginning to match.
Of course, for more prosaic needs, focusing on vectorizing key operations can solve speed troubles. And it's worth noting that the $1,000+ per copy that Stata costs can buy an awful lot of extra processing power to throw at the problem.
Fortunately, Radford Neal has been hard at work speeding R up, and has released some new patches to play with if you find it too slow. You can also try writing key sections in C++, or using Revolution Analytics' offerings (free for academics).
For extreme speed needs, however, R can't be beat, as it has long offered graphics-card based extreme parallelism that commercial solutions are only beginning to match.
Of course, for more prosaic needs, focusing on vectorizing key operations can solve speed troubles. And it's worth noting that the $1,000+ per copy that Stata costs can buy an awful lot of extra processing power to throw at the problem.
Thursday, March 10, 2011
R: Drop factor levels in a dataset
R has factors, which are very cool (and somewhat analogous to labeled levels in Stata). Unfortunately, the factor list sticks around even if you remove some data such that no examples of a particular level still exist
# Create some fake data
x <- as.factor(sample(head(colors()),100,replace=TRUE))
levels(x)
x <- x[x!="aliceblue"]
levels(x) # still the same levels
table(x) # even though one level has 0 entries!
The solution is simple: run factor() again:
x <- factor(x)
levels(x)
If you need to do this on many factors at once (as is the case with a data.frame containing several columns of factors), use drop.levels() from the gdata package:
x <- x[x!="antiquewhite1"]
df <- data.frame(a=x,b=x,c=x)
df <- drop.levels(df)
Now I'm going to quit monkeying around and get to sleep.
# Create some fake data
x <- as.factor(sample(head(colors()),100,replace=TRUE))
levels(x)
x <- x[x!="aliceblue"]
levels(x) # still the same levels
table(x) # even though one level has 0 entries!
The solution is simple: run factor() again:
x <- factor(x)
levels(x)
If you need to do this on many factors at once (as is the case with a data.frame containing several columns of factors), use drop.levels() from the gdata package:
x <- x[x!="antiquewhite1"]
df <- data.frame(a=x,b=x,c=x)
df <- drop.levels(df)
Now I'm going to quit monkeying around and get to sleep.
Monday, February 7, 2011
R: Functions and environments and a debugging utility, oh my!
The Mark Fredrickson blog has a superb post on R functions and environments that's well worth checking out.
He also includes a handy function for debugging:
> fnpeek <- function(f, name = NULL) {
+ env <- environment(f)
+ if (is.null(name)) {
+ return(ls(envir = env))
+ }
+ if (name %in% ls(envir = env)) {
+ return(get(name, env))
+ }
+ return(NULL)
+ }
> fnpeek(f1)
[1] "n"
> fnpeek(f1, "n")
[1] 7
He also includes a handy function for debugging:
> fnpeek <- function(f, name = NULL) {
+ env <- environment(f)
+ if (is.null(name)) {
+ return(ls(envir = env))
+ }
+ if (name %in% ls(envir = env)) {
+ return(get(name, env))
+ }
+ return(NULL)
+ }
> fnpeek(f1)
[1] "n"
> fnpeek(f1, "n")
[1] 7
Friday, February 4, 2011
R: Colors
"It ain't easy being green. ~ Kermit T. F.
Matt Blackwell at the SSSB has made it easy to access all the Craylola(tm) colors in R.
And in case you're not familiar with the way R handles color, here are a few resources:
* The best color chart for R.
* Color palettes in R (allows plotting a spectrum or coordinated palette of colors easily).
Matt Blackwell at the SSSB has made it easy to access all the Craylola(tm) colors in R.
And in case you're not familiar with the way R handles color, here are a few resources:
* The best color chart for R.
* Color palettes in R (allows plotting a spectrum or coordinated palette of colors easily).
Thursday, February 3, 2011
Regular expressions, an example
Why regular expressions are your friend. This is written in Stata but applies to any language where regular expressions exist.
Original version
if length("`qx'")==3 { /*ex: 5q0*/
local a = substr("`qx'", 1, 1)
local b = substr("`qx'", 3, 1)
}
else if substr("`qx'", 2, 1) == "q" & length("`qx'")==4 { /*ex: 5q20*/
local a = substr("`qx'", 1, 1)
local b = substr("`qx'", 3, 2)
}
else if substr("`qx'", 3, 1) == "q" & length("`qx'")==4 { /*ex: 10q5*/
local a = substr("`qx'", 1, 2)
local b = substr("`qx'", 4, 1)
}
else if length("`qx'")==5 { /*ex: 10q20*/
local a = substr("`qx'", 1, 2)
local b = substr("`qx'", 4, 2)
}
Regular expression version
foreach q in `qx' {
local a = regexr("`q'","q.+","")
local b = regexr("`q'",".+q","")
}
Original version
if length("`qx'")==3 { /*ex: 5q0*/
local a = substr("`qx'", 1, 1)
local b = substr("`qx'", 3, 1)
}
else if substr("`qx'", 2, 1) == "q" & length("`qx'")==4 { /*ex: 5q20*/
local a = substr("`qx'", 1, 1)
local b = substr("`qx'", 3, 2)
}
else if substr("`qx'", 3, 1) == "q" & length("`qx'")==4 { /*ex: 10q5*/
local a = substr("`qx'", 1, 2)
local b = substr("`qx'", 4, 1)
}
else if length("`qx'")==5 { /*ex: 10q20*/
local a = substr("`qx'", 1, 2)
local b = substr("`qx'", 4, 2)
}
Regular expression version
foreach q in `qx' {
local a = regexr("`q'","q.+","")
local b = regexr("`q'",".+q","")
}
Sunday, January 30, 2011
Tab completion
Let's say your hands are aching from too much typing in of variables. What to do? Get a keyboard tray and learn proper ergonomics, of course.
But what if you just want to reduce the amount of typing in of variables you do for reasons of laziness...err...efficiency. Well, you can type the part of the variable that's unique and then hit Tab. Stata or R (or many other programming environments) will both fill in the rest for you.
Suppose you have variables named:
AriWillGetRejectedFromHisFavoriteSchools
MaraWillGetInEverywhere
MarkWhosMark
DumplingsAndData
If you wanted the first variable to show up. Just type "A" and hit TAB.
But if you want the second, you'd have to type "Mara" and hit TAB, because until you hit the fourth letter it won't be sure which variable you want.
But what if you just want to reduce the amount of typing in of variables you do for reasons of laziness...err...efficiency. Well, you can type the part of the variable that's unique and then hit Tab. Stata or R (or many other programming environments) will both fill in the rest for you.
Suppose you have variables named:
AriWillGetRejectedFromHisFavoriteSchools
MaraWillGetInEverywhere
MarkWhosMark
DumplingsAndData
If you wanted the first variable to show up. Just type "A" and hit TAB.
But if you want the second, you'd have to type "Mara" and hit TAB, because until you hit the fourth letter it won't be sure which variable you want.
Sunday, January 23, 2011
STATA: Regular expressions
A regular expression allows you to do a moderately fancy search (and replace if you want). So say you wanted to replace all the "Dennis"s in a variable with "Awesome"s, but only if they're at the end of the line. You could try:
-replace PBFnamevar = regexr(PBFnamevar,"Dennis$","Awesome")-
You could also replace any character, or just capitals, or just digits...there are lots of possibilities:
http://www.stata.com/support/faqs/data/regex.html
You can also use it for locals:
-local strata = regexr("agecat","age")-
Or -if- commands:
if regexm("`strata'","age") {
}
On a related note (although not actually regular expressions), say that you've got a string variable that consists of a bunch of what should be separate variables, only lumped all into one, separated by a semicolon (e.g. a row might look like "1;15.2;89;hi;21"). Try -split-:
-split textvar, gen(newtextvars) parse(";")-
I should note that Stata's regular expressions are wimpy compared to what other languages support. R supports PERL regular expressions, which can do so many things it's scary.
-replace PBFnamevar = regexr(PBFnamevar,"Dennis$","Awesome")-
You could also replace any character, or just capitals, or just digits...there are lots of possibilities:
http://www.stata.com/support/faqs/data/regex.html
You can also use it for locals:
-local strata = regexr("agecat","age")-
Or -if- commands:
if regexm("`strata'","age") {
}
On a related note (although not actually regular expressions), say that you've got a string variable that consists of a bunch of what should be separate variables, only lumped all into one, separated by a semicolon (e.g. a row might look like "1;15.2;89;hi;21"). Try -split-:
-split textvar, gen(newtextvars) parse(";")-
I should note that Stata's regular expressions are wimpy compared to what other languages support. R supports PERL regular expressions, which can do so many things it's scary.
Subscribe to:
Posts (Atom)