Posts tonen met het label SAS. Alle posts tonen
Posts tonen met het label SAS. Alle posts tonen

woensdag 5 september 2012

Cornell on statistical packages and Ethan Fosse's blog

This guy, Ethan Fosse, is a sociologist who uses R. Instead of Stata. Picture this. He does know a lot about graphics but I particularly liked this corner of the blog where he points to unsolved questions in sociology. Rising inequality in any field is one such question that intriges me too. Interesting readings. Found it through this R blog (Revolutions), when I was actually looking for sociologist using Stata. And yes, it's what people at Cornell recommend to learn first. I quote:


General Purpose Software

Quantitative data analysis in sociology is dominated by three all-purpose programs, Stata, SPSS, and SAS. All three are available on Cornell's Athena computer cluster, and all three are excellent.
We typically recommend that graduate students learn Stata first. SPSS has some attractive features, but it has been losing ground to the other two for the past 15 years. Although SAS is the most comprehensive, it has a relatively inefficient programming language. Stata is almost as comprehensive as SAS, and it has a much more efficient programming language. And, if you need one of the special routines that SAS offers but Stata does not, you will often end up using a more specialized computer program anyway because even SAS is not quite as good as the specialized software (see below). Nonetheless, SAS is particularly well-suited for very large datasets, and SAS is also the dominant package for the federal government. If you work with many types of government data, you will find that the best supporting documentation is written for SAS users.
UCLA statistical computing has the best (we think) on-line set of resources for these three programs. See http://www.ats.ucla.edu/stat/. Also, Cornell's CISER offers tutorials for all three programs. Seehttp://www.ciser.cornell.edu/ASPs/workshops.aspx.



donderdag 29 juli 2010

What stats package to use?

Introduction

Boys like their toys, and this is not different with statistical packages. It's a perpetual and heated debate and when you've landed at some point and think your workflow is good, technology passes you and sets you back. 

Here's an old discussion that I first consulted, but below I make my own considerations. In short: I'd use Python for big data, Stata for analysis, and R if I have to (e.g. for some graphs). Everything else I would ditch.

Stata

Stata is my program of choice. It is quite expensive, but mind you that for a couple of hundred euros, depending on the flavour, you'll not only get an easy and robust statistical software package, but also fast support, a good community, useful user commands, and a great documentation source. Fun fact: all documentation is read by the wife of the founder, who's not a statistician but perhaps even smarter. If she doesn't understand what the statisticians are saying, it goes back to the drawing board. 

The bad things: forget about ever copy-pasting anything. You'll also need to have a lot of memory on your computer, as Stata loads the whole file and just one at a time (although you can 'preserve' a file temporarily to work on something else in between).

Python

Python is the next language I will learn. I have used chunks without understanding what I was doing, but I like the sound of the language, and it's the logical step-up after Stata, it seems. Many people are using it and so will I.

R

I don't like R. There is a thorough discussion here, circling around leaving Stata for R, but ending up in concluding what I conclude about: R is a mixture of a coding language like Python and a statistical language like Stata, but because it is open source the support is unsure, the community tends to be geeky and unfriendly, the documentation is poor, and the language consistency - even if the structure is good because it is a programming language - is bad. Some commands have their own inner programming language and that is plain bad. 

The good things: it is free and R Studio is a great user interface. It has good graphic capabilities, and 

Mplus

I don't know Mplus. Colleagues use it when there are issues with missing values, and the programmers are said to be the best statisticians in the world. So it must be good, but I don't use it.

SAS

This is old software. It is too complicated, and while it can do a lot through obscure options, it is not flexible enough to do what you want.

SPSS

This is bad software. It is a scandal that some universities still teach this.

Some R resources

Apparently the single best manual for R: https://r4ds.had.co.nz.

dinsdag 22 december 2009

Start to Stata

Ik was een SPSS gebruiker, om twee redenen:
- Dit is wat men aan de universiteit aanleerde
- SAS is een lelijk beestje

De meeste onderzoekers zullen erkennen dat de mogelijkheden van elk pakket hen boven het hoofd gaan, ook van het 'speelgoed' SPSS. Maar afhankelijk van de taak die je moet uitvoeren kan je een voorkeur hebben voor bepaalde software. Omdat ik gek werd van de syntaxcontrole in SPSS, en bepaalde econometrische tests niet vond, probeer ik nu Stata uit.

First impressions of Stata

Eerste opmerking: het categorizeren van variabelen (prefix i) en inbouwen van interactie-effecten (# en ##) is geniaal.

Eerste klacht: grote datasets raken niet geladen. Mijn computer heeft 3 GB ram, maar toch kan ik slechts 700 MB aan stata toewijzen. Dit is vreemd en vervelend, aangezien sommige administratieve data die ik gebruik groter zijn dan 1 GB. Toewijzen gebeurt als volgt:

set memory 700m, permanently

Jammer genoeg moet ik de data dus eerst in SPSS laden en opsplitsen tot ze bruikbaar zijn in STATA.

Coming from SAS?

Here are some websites with syntax translation:

Introductions