Posts tonen met het label English. Alle posts tonen
Posts tonen met het label English. Alle posts tonen

vrijdag 1 februari 2019

Billionaires

It seems that in the world today, the 1% are escaping their responsibilities. They cannot get away with only philanthropy – billionaires should pay "a fair share" of taxes, which means higher marginal taxes on wealth. At the Davos summit, Rutger Bregman referred to the 91% income tax under Eisenhower.

Others will counter that the Laffer curve needs to be taken into account: at some marginal tax level, tax avoidance will become an issue. Sure, but we also got technology to counter this issue that Laffer didn't have at his disposal.

So naturally I'm in favour of taxation, for the simple reason that income inequality should be limited in order to maintain social equality and an inclusive economy to which each can contribute and in which each can consume. This is even a condition for growth.

Yet I don't agree with the 1% rhetoric. Instead, I believe that as organizations, enterprises, and people acquire wealth over certain thresholds, their influence expands for good reasons. The market is also some form of democracy, and as with other forms, like the electoral meritocracy, it has its pitfalls. However, with great power comes great responsibility. The accountability of the rich should go beyond paying taxes. They have an exemplary function and should contribute to society in a positive way, and this is a matter of morals but also of monitoring. The 1% should work together with the state to provide for the 99%.

Then I don't see a clear distinction between a national bank that takes money out of the economy to increase interest rates, and a single person or organisation that does the same. Then I don't see a difference between schooling projects funded by a government agency, or by a committee nominated by the 1%. Of course, it is not just that a very small number of people are a millionfold better off than a very large number of people, and when I say that they should provide for them, I mean that they should. In the same way, the people can accept or overthrow a king. The 1% should deliver or go down. What you pay to them is just an additional tax, and you should expect value for money, just as you expect from the State.

The real issue is uncontrollable wealth that distorts social life: an elite of 10% to 20% that enjoys the pleasures of this world based on exploitation of the ones below. They should be taxed more, because we expect the morals, but less the actions that we expect from the 1%. Let's call this group the 19%, and let them not believe they are the 1% in the making. They will never be if they don't occupy a natural monopolist position, as Google, Facebook, Microsoft in the digital world do, and a limited number of car manufacturers and petrochemical companies in the manufacturing and utility sector. The latter are identifiable, the former are anonymous but perhaps more pretentious. We can just take the yacht.

donderdag 25 oktober 2018

Four conditions for social dialogue

What makes collective bargaining work and what makes it worthy of institutional support? In a recent EU workshop, we have been discussing the case of Belgian and also possible lessons for other EU Member States. I will briefly summarize the discussion and draw my personal conclusions.

The social dialogue construction

I like to compare social dialogue with a simple building that has four walls and a roof. Most people will agree that a roof is almost sufficient as a shelter, but it needs the walls to support its function. Note that at the heart of social dialogue, which is bipartite (between employers' federations and trade unions) or tripartite (adding the state or government), there is always some form of collective bargaining. The terms are therefore often interchanged.

Coincidentally, we came up with four conditions for a solid structure for social dialgue:

1. Representativeness

This answers the question of who is sitting around the table. The legitimacy of social dialogue depends on the position of the participating actors. For instance, the unionization rate should be high enough to show the force of collective action, while the employers' density rate equally should be high to ensure enforcement of the agreements and no undercutting by rogue employers. It could be that the actors involve only represent a part of the economy or a limited number of companies.

2. Rules and procedures

Social dialogue can be an informal or a formal matter, but it should ideally take place within a legal or paralegal framework that sets out the structure, procedures, and rules. This should be clear, transparent, and with some degree of universality. In closed shop negotiations, the rule is that what is negotiated, only affect affiliated parties, and no other parties (e.g. non-unionized workers) are allowed. More developped collective bargaining has a much wider scope, leading to agreements that are binding to all parties, even non-affiliated, through legal extension. This means that the agreements, which essentially are private contracts, obtain the status of a law. This is the material aspect of social dialogue, answering the question how dialogue is done.

3. The process

The rules and procedures are an empty box that needs to be filled. This happens through a process of negotiation, and the contents answer the question what is the stake in the game. This is rather intangible, and relies to the influence the parties have, the sense of ownership and responsibility, the shared goals and trust in the system, and the ability to reach a consensus or compromise. Beside the fact that collective bargaining is essentially a power game, social dialogue also relies on a mutual understanding.

4. The outcomes

Social dialogue is sometimes an end in itself, ensuring information and communication about decisions that are taken by the organization, and allowing some codetermination. It may also reach further, contributing to the interests of an economic sector or the national economy, as well as other social goals (inequality, poverty, democracy). Social dialogue is justifying its existence by answering the question why it takes place. This also implies a cost-benefit calculation: as negotiations take time and take away individual degrees of freedom, which are replaced by collective arrangements, the benefits should compensate the losses. An easy example is an agreement that would need to be negotiated in every company and lead to similar results, is better done just once. Also, some agreements would be preferred by most companies if all others would do the same, but avoided if there was no guarantee that this would take place. For instance, higher salaries may respond to the reservation wage of workers, but a firm that implements the higher wage would be outpricing itself on the market, while others continue a practice that leads to a loss in the workforce and therefore for the sector, but at least manage to survive another turn.

Lessons for Belgium

The Belgian system of collective bargaining is a well-known archetypical example of post WW II social dialogue. It is the end of a buildup that actually started after WW I, where the working class organized its interests against the power of the ruling capitalist class. It is also a compromise between two ideologies: on the one hand a militant socialist movement that strives for a revolution, and on the other hand the corporatist, Christian, anti-socialist and even fascist idea of a common interest between workers and employers. It is a little known fact that indeed the nazi's spurred employee participation of employees in the workplace during the occupation of the country. Yet both the elements of opposition and interdependence are present as it is. For instance, collective action (strikes) is less constrained as in other countries, and trade union representatives are protected against dismissal, while at the same time, trade unions are represented in the OSH committee, in the company council (but without decision power), in sectoral and national joint committees and tripartite bodies.

An interesting observation is that the 'outside options' currently go in two directions. On the one hand, there is popular protest and the threat of striking, on the other hand, there is the state that represents the interests of the ruling class. As a side note, this is a critique of the state that liberals and socialists share, although the former are accused of being 'neoliberal' (e.g. effectively using the state to manipulate the free market), and the latter of being 'statist' (e.g. aiming for a maximal state), which is twice correct.

Either way, although social dialogue / collective bargaining is very well organized and institutionally integrated in Belgium, it's independence is not absolute. Even if it never was, and some 'grease' to facilitate bargaining has often been necessary, today this takes a legal form that explicitly constrains collective bargaining: because of a lack of trust by the 'third party', the state imposes a number of laws and procedures (e.g. on wage growth, training plans, ageing plans, gender inequality. To the outsider, or to the social partner that benefits from the state option, even if it is an agreement of last resort (in which case the negotiator only loses time by not reaching an agreement in his favour), this may seem wise. In reality, it is a loss of sovereignty and independence of collective bargaining, taking away the benefits of the system and mainly its flexibility, as the government is only a remote player. However, it is not only the lack of trust by the government and its ambition to plan the economy that reduces the scope of collective bargaining, but also the belief of the social partners themselves that collective bargaining will deliver the desired results or that a consensus or compromise is possible. Hence there is a couple of bricks missing in the third wall (the process), as they have been used to fortify the second (rules and procedures). The roof is still on the building, but it is more unstable now.

Lessons for Europe

The problem of best practices is that there is a reason why institutional structures have a certain form, and therefore a nice dress doesn't always fit on someone else. This is the functionalist view but it should not take away the ambition to converge. Yet it should be clear that the practices of Belgium, Germany, and the nordic countries cannot be immediately introduced in, for instance, the new members states with a much weaker tradition of social dialogue. Even if all parties would agree that the building consists of those four walls and a roof, and then allows social and economic development, there are no incentives to start building it if there are no immediate gains.

This perspective is most often totally lacking in the political discourse: the reason why workers join a trade union is not macro-economic growth. A union has three functional levels: the first, as a service provider; the second, as a representative organization; and the third, as a social actor. The first level, services, is what matters most: trade unions provide legal protection, voice at the workplace, and in the case of Belgium, do the administration of unemployment benefits, and aid in the shifts from and towards sickness leave, training, and retirement. The second level is, representation, is more important from a legal and institutional point of view - concluding agreements - but in countries where this is all that remains of workers' organization, it quickly fades out because of free-riding tendencies. The third level, social action, is ideological, and relates in part to the origins of the trade movements: one from the revolutionary socialist movement, one from the corporatist christian movement, and one from the liberal movement that recognizes the freedom to organize interests. At the third level, trade unions also contribute to democracy,  as the main player in the civil society, with a more direct link with the members in contrast to political parties that are periodically elected and where no membership fee is required. Hence, the vote that workers cast is 'voting with the feet', a saying by Lenin that was taken up by Tiebout and many scholars (Hayek, Friedman, Hirschman) later on.

Hence we should be careful on the one hand not to confuse the origins and the structure of the system (the Althusser argument). Other countries will not replay the history of Belgium, nor will it be possible to directly copy the structure, even if they have the finished building in mind. However, too often the statements of either party, whether it is the state, the trade unions, or business organizations, is that institutional traditions are too different and welfare state levels cannot be generalized, so that nothing should be undertaken. This goes beyond the fact that good ideas don't stop at the border and beyond the need for higher-level coordination of social developments in a unified economical sphere. Rather, a number of steps can be set to build up the four walls, so that already they provide shelter against the wind, if not against the rain. For instance: workers can gain influence at the company level, and discussions on OSH can lead to more safety as a direct outcome that is more efficient than legal formalism. Rules and procedures are developed on the go, and representativeness is an endogenous outcome when gradually more rights are given to the parties. In other words: the stepping stones are building bricks. The role of the state is rather to nudge and facilitate than to impose practices that lack foundation. This is not only the case in developing systems of social dialogue, but also, as outlined above, in mature regimes.

vrijdag 14 september 2018

The labour share and disemployment

A Spanish unionist asked me what happens if productivity increases because of disemployment. Would the 'wage rule' still hold? The wage rule is that wages should follow productivity and inflation, in order to keep the wage share constant.

The answer is yes, because the composition of the work force and output have both changed. Suppose low-paid, low-productivity workers were dismissed, then the output per worker will increase, but as the high-productivity workers are already better paid, this implies a natural increase of the average wage. You don't have to negotiate anything. However, if there is an additional increase in productivity, because of productivity enhancing investments, that would be the basis for wage claims.

This choice implies that you stick with the 'single productivity' measure. Wage claims in one, suppose a labour-intensive, sector based on overal productivity growth will cause the labour share in that sector to further increase. Now, on the demand side it does not matter which sector delivers the disposable income, but at the labour market, deviations from the sectoral labour share derive from the production technology (the amount of capital and labour that is optimal for any production level). This implies that a shift in the structure of the economy will create imbalances: suppose well-paying sectors would entirely disappear and only the car industry remains with high output but low wages, then there would not be enough disposable income and either car prices would drop or wages should increase.

woensdag 21 maart 2018

Composition bias in small samples with skewed distributions

If regressing on wages that are not normally distributed (right-skewed, in general), and comparing groups with very different relative frequencies (like men and women in the STEM-field), added a censoring (e.g. wage floors), then a positive bias for the larger group is likely for small sample sizes.

I demonstrate this below using a simple Monte Carlo simulation.

In a tobit-like fashion, one would like to control for the selection chance, but this is given by gender which is already in the model.

It would be interesting and easy to make this simulation based on actual means, sd, skewness in the (male) population.


clear
cap erase mc.dta
foreach n of numlist 60 80 100 120 140 200 300 500 800 1000 2000 {
forvalues r = 1/500 {
clear
*local n = 1000
set obs `n'
local m = 2000 // mean
local s = 200 // standard deviation (spread)
local f = 0.1 // feminisation
local a = 3 // positive: right skewed (it is possible to compute the skewness metric based on a)
gene g = 1-(runiform()<`f')
gene rn = runiform()
gene d = 2*invnormal(rn)*normal(rn*`a')
gene w = max(1400,`m'+d*`s')
*twoway kdensity w
collapse w, by(g)
gene nsize = `n'
gene c = 1
reshape wide w, i(c) j(g)
list 
cap gene gpg = w1/w0
cap gene gpg = .
cap append using mc.dta
cap save mc.dta, replace
}
}
replace gpg = round(100*gpg-100,.01)
tabstat gpg, by(nsize)
*scatter gpg nsize

exit

donderdag 5 januari 2017

Football inflation

The wages of football players are astonishingly high and we do not precisely know why. In this tentative post I explore a few possible explanations. The main argument is that we have a market inefficiency and it is built on a simple intuition: you do not have to pay unschooled labour very much to make them hit a ball all day. Yes, the talent is rare, and the very best will not refuse more money if you throw it at them. And yes, you will loose the bidding process if you don't show the money, so the exchange value of elite football players may be very high, but no, there are no efficiencies gained during the wage escalation that justify the wages.

The standard way of thinking is that it is all propensity to pay, and due to a superstar effect of millions of people watching just one team that has the world's most famous player, there is enough at stake to raise wages sky-high. However, what profits does such a team make? Not a lot, as there is a finite capacity in the stadion, and you cannot ever have enough markups on some letters printed on a shirt. Let's discard that road. Broadcasting rights then, yes, amount to many millions that are injected in the world of sports. Again, if the broadcasting rights were cheapish, or non-existing, would that be a game-changer? My guess is no: I was watching all games in any league in the 80s when broadcasting rights were negligible, and players were equally motivated as today. The main difference is that back then there were no ads during the games, so we actually got more sports for less money, in fact for free. Broadcasting rights are very similar to the author's rights in the music and movie industry. They are actually bad calculations. To have your actors, players, musicians, show up and do all the necessary recording, training, and coaching, a known sum of money has to be paid. You can then figure the audience and the expected returns at any price level. At some equilibrium price level, welfare is optimal, and a monopolist will go above that equilibrium, restricting the quantity and increasing prices. But what is more: long after all wages and input costs have been paid, the intellectual (well...) rights still generate returns. As a consumer, you get no value for your money, because production is at zero marginal costs. Now it could be me, but paying more for nothing extra seems very ... inflationary.

What is worse is that the price effects may be endogenous. Say the market for sports products is not very competitive, a paradox indeed, but probably true. Marketing defines what brands are in the market, and multiple brands have the same shareholder structure. I'm no specialist in the field, but I read for instance that Puma and Adidas come from the same family business, and they largely avoid competition, with Puma focussing on the expensive stylish segment of the market, and Adidas on more specialized sportswear. Suppose also that there is a fairly inelastic demands for those fashionable goods, in fact the goods from the companies that decide what's fashionable. Then they can bid whatever amount of money to sponsor the teams that pay the players, and likewise, any other sponsor can further inflate prices by sponsoring the television channels for the broadcasting rights. It is an orgy of inelastic demand and price raises. Interestingly, in Belgium there are many ads for cheese and ham with a local label of origin. By law, these companies have no competitors in the same market. So these companies cause inflation, with a peculiar consequence: inflation leaves room for inequality. The first to benefit are the players, of course, and the last are the spectators that need to ask for a pay rise to their bosses. With some luck, when working at one of the sponsoring companies, that is not too hard, but others will lag behind.

If this sounds like a bubble, that is what it is. It is remarkable that organizations that pay extraordinary salaries, like football clubs, have equally high debts. This means the banks recognize the value of the players as real assets. This is foolish, as everyone know that one tackle suffices to end a player's career, but probably they insure against this. Anyway, it certainly is economically dangerous and also suspicious. A lot of money is flowing in from Arabic states, from oil sheiks and the like. Yes, they produce inelastic goods, but they also excel in money laundering, as do the teams they employ.

In sum, football is inflationary, very literally: thin air blowing up a bubble.

dinsdag 29 november 2016

Execute R code in Stata / Read SPSS data files

On 64-bit Windows OS and on Mac, the -usespss- command does not work. If you want to use SPSS data (or SAS), without quitting Stata, do something like this:

rsource, terminator(END_OF_R) rpath(R_pathname)
library(foreign);
rprecar<-read.spss("precar.sav", convert.f=TRUE);
rprecar
attributes(rprecar);
write.csv(rprecar,"rprecar.csv", na = ".");
q();
END_OF_R

It may be tricky to find your R_pathname. I didn't bother looking up what it is for now - probably the path to the executable (which sucks, because on every computer it will be different). In the example the conversion goes to csv, and there's the missing values ("na") option. The foreign package actually also allows old Stata 9 files, which is just fine and will preserve most labels, variable names, and missing values.

Here's the info for -foreign-:
https://cran.r-project.org/web/packages/foreign/foreign.pdf

Taxing robots or lowering minimum wages?

There's some errors in the logic, don't believe this.

There is a lot of talk nowadays in Belgium on either taxing robots or lowering the minimum wages. Remarkably, typically right-wing parties oppose the first and opt for the second, while typically left-wing parties propose the first and object to the latter. Personaly, and consistently, I am not in favour of any of the two measures, because indeed... they are the same.

Basically, by taxing robots, capital is made more expensive, which causes a shift to labour as an input factor. Similarly, the relative price of labour may go down by lowering the minimum wage, causing the exact same shift between the two input factors, all else being equal.

As I explained in another post in Dutch, one may argue whether to tax capital in general (profits, wealth, etc.) or wages. There could be an equivalence between all tax options. The main worry is that taxes cannot be avoided, which is why it is generally preferable to tax in tiny bits everywhere (capital, consumption, income, transfers, real estate, inheritance, etc.) - this is my theory of chaos economics. However, for now it is clear that labour is easily traceable, and hence an easy target for taxes. The only condition is that wages are sufficiently high, so what matters is not the tax rate - which you should calculate over all sources - but the wage share, which should be around 75% for a steadily growing economy, if the second half of the 20th century is to be taken as an example.

So please, do not tax robots - increase wages!

zondag 6 november 2016

Zero marginal costs

There is an interesting phenomenon on its way which we do not fully understand yet. It's the computer revolution. Now many observer believe that the second machine age is like the first, that technology has always been seen as a threat but in hindsight created more work and welfare. I do not buy that for a number of reasons.

First, the simple stat is flawed. Yes, more people work and there is more welfare, but there are also many more people living on earth compared with the middle ages. We do not have precise observations, but it is not unlikely that the majority of the people before the industrial revolution were all working full time, regardless of age and sex. So the machines may have enhanced the labour force, and freed up time for leisure for most, but some are unlucky not to have the skills that suit a machine. In order to let them accept the introduction of technology, we should therefore always distribute welfare, also to those who have made room for progress.

Second, in this machine age, I believe we are heading towards a point were the very human aspect of labour, not its force, but its knowledge, is being substituted by computers. That is very different, because it excludes complementarity of capital and labour to a much larger extent. For now, this leads to job polarization, but soon software may take over more advanced tasks which have now fairly big gains, and robots might take over many service jobs that require a 'human touch'.

Third, the apologists say fear is not needed as the past predicts the future and we've sorted it out last time. I believe that is the wrong motivation. Fear is not needed, indeed, but not because nothing is going to change, but because it is possible to see a change for the better. Why, after all, would we want to work? We don't, that's why we ask a wage in return for the trade of leisure. For those who think this will lead to a demographic boom, maybe not: if we spend more time caring about individuals, we have a new constraint other than the price of education: the time which we cannot multiply. Also, procreation brings responsabilities, which we may want to keep under control.

Fourth, I tend to agree with Paul Mason that there will be a change of system. The steam machine introduced capitalism. After all, those machine are capital that allows the making of profits, of which part goes to the worker and the rest to the shareholder. The new technology works autonomous, which will actually take away profit. We tend to think of companies that adopt technology as labour-killers, but what they really are is profit-killers. The beauty of capitalism is that companies will always want to be the first to go on that road, because then profits actually increase. As soon as one more company follows, we have Bertrand competition that drives prices to zero, because software (the new capital) produces its goods and services at zero marginal costs. I thought of this a while back, and others have without a doubt too, but Paul Mason found a really interesting note of Karl Marx who predicted exactly this, and it is entirely in line with marxist economics. If labour becomes obsolete, there is no more income and no more turnover. So it is a system changer. Now before we yahoo about this, nobody knows what the new system will be. It is possible that we distribute according to needs, or based on a basic income scheme equally for all. It is also possible that capital holders will manage to protect intellectual right on software and intangible products (like a song or a movie), and have ordinary people sweat to obtain it without actually contributing much to society. Amongst capitalists, there will then be a fair amount of trade, owning brands and intellectual rights, and forming cartels much like today, while those who were former workers will then be slaves. Personally I think the latter scenario is most likely, even if it relies on reducing democracy. It is probably even already ongoing.

So in sum, there will be a change and it could be for the better or worse, depending on whether you catalogue more equality as the former or the latter. It certainly deserves a debate.

maandag 21 september 2015

Strongly balanced sample (Stata simulation)

clear
local base = 4      /* Set nr. of observations per units */
local size = 200    /* Set nr. of units */

*****
local tot = `size' + `base' - 1
di `tot'
set obs `tot'
gene id = floor(_n / `base')
gene time = `base'*(_n/`base' - id) + 1
drop if _n < `base'

*****
xtset id time

dinsdag 7 juli 2015

Using -reindex- to calculate real trends

We live in a real world, not in a nominal fiction. Hence it is practical and sensible to express amounts of money in real terms. I will show how to use my Stata user command - reindex -  to help you with this.

There is no straightforward way of deflating. When the monetary base expands, inflation may result. But it is also possible that economic activity increases, in which case you can still buy one can of coke with one euro, but as you happen to have two euros, you can buy two.

There are four standard ways of deflating prices:

  1. The GDP deflator is the ratio between the nominal gdp and the 'real' gdp which is the counterfactual with prices of the reference year. Even if it seems a logical way of calculating price increases, there are numerous biases in this measure - I can't find out what to do with new goods that weren't priced the year before. The standard GDP deflator is measured based on goods produced in the economy. 
  2. The deflator of consumed goods is exactly the same but with the current goods that are purchased. Importantly, the basket of goods changes every year.
  3. The (national) consumer price index is a fixed basket of goods corresponding to the average household's consumption, of which prices are tracked over time. Only in the longer term the basket is changed, often when goods have become out of fashion for a while already. Sometimes there are minor adjustments to the weights of goods.
  4. The harmonised consumer price index is the same as above but with an internationally harmonised basket of goods. Of all deflators I find this one capturing inflation best.
Data may be monthly, quarterly, come as an index or as a yearly change. Before continuing, make sure you transform into the time period of the nominal variable you want to convert. Indicate the unit within i() and the year variable in j(). Merge the data based on year and units.

Now use -reindex- to set both the deflator and the nominal variable to the same type. Reindex can work with four types:
  • Levels
  • Index
  • Percentages
  • Factors
Note that levels can only be converted into the trend types, not the other way around. Note also that percentages and factors are actually the same thing, but the percentage change is the factor change minus one.

For percentage trends, we commonly find an expression like the following:

%.real growth = %.nominal growth - %.inflation

This is a rule of thumb that is ok for small numbers. However, the right calculation is in factors:

nominal growth / real growth = inflation,

which in logs would be ln.nominal - ln.real = ln.inflation and as the log of a factor between .90 and 1.10 is very close to the percentual change (i.e. factor - 1), the 'easy' formula above would hold.

We can move factors to obtain:

real = nominal / inflation

The above we can also use with indices, where you set the base for all indices in one and the same year marked by tobase() to 100 using toscale(), except for the deflator (inflation) which you set to 1. The reindex program allows setting the base time and scale. 

donderdag 30 april 2015

Markdown

I'm a sucker for simplicity. I found Word too multifunctional, I found LaTeX too hard. Then Markdown appeared, the language used on Wikipedia, on my Trello app, and - commonly - in emails.

Markdown syntax

There are some complication, like hyperlinks that I don't want to mention here, in order not to obscure the absolute simplicity of Markdown. This is what you'll use:
# Heading 1
## Heading 2
### Heading 3 (you get the point)
*bold*
**italic**
![Graph alt text](./path.png) 
Nothing more.

Markdown on OS X

One address: MacDown. It is a beautiful program that has a split screen set-up: left you have a WYSIWYM interface, right there is a live preview. Did I say it is beautiful? It's free.

Markdown on Windows

I have tried a few, but on Windows with no luck. There is always some premium function for sale and I don't like that. WriteMonkey is a good idea, but the pdf support is ... not included. There are also a bunch of web apps, most notably Dillinger and StackEdit. They look good, but I like to open a file from windows explorer and then that's just a pain.

In fact, gVim will do all you want, but there's a learning curve as it doesn't behave like normal textpads (I like notepad++ for instance as an in-between solution).

Two programs however will do the job, and they are software packages! Obviously, R is free, and Stata is not. You will know that I love Stata - I think it's the best, albeit expensive, textbook in statistics and it also has free software with it.

Markdown in R and Rstudio

Rstudio is brilliant software. Use the -knitr- package to enjoy writing markdown files and exporting to pdf. It's made for that and it's used by millions.

Markdown in Stata

Stata is a different beast. In a way it is less flexible and the markdown integration doesn't feel as native as in R. However, it is also easier.

You output an smlc log file, then you parse the file with the user command -markdoc- (Converting SMCL to Markdown and other formats using Pandoc). That's all there is to it. Remember to put your markdown code between command /* beginnings and endings */ which you put on separate lines. Also use - quietly - to reduce output.

There are a couple of other Stata programs doing a similar job:
  • Weaver: HTML and PDF dynamic report producer
  • Ketchup: HTML and PDF dynamic report producer
  • Synlight: SMCL to HTML convertor and syntax highlighter

woensdag 29 april 2015

Weights in Stata

Stata has four weights, your average statistical software has one. Still Stata is right, as I will try to explain in even simpler words than elsewhere. But let's ignore the iweight for programmers, and focus on the other three:

fweight or frequency weight - is probably the easiest, but most abused. It says that one observation represents the number indicated by the weight. Imagine you collapse a dataset based on gender, region, and educational attainment, and you regress education on gender, then the count of each line would be your fweight. Data is commonly stored in this way to reduce duplicate lines. It follows that fweights should be integers, because there is no such thing as a half respons.

pweight or sampling weight. For instance, you need to have 20% of women with an academic degree, but for some reason the sampling only gathered 10%, so you will want to double their weight. Using fweigh would overestimate the number of cases but underestimate the variance. In other software, you might rescale the weight so that it sums to the original n, but using pweight is better. Stata leaves you no choice because fweight does not work with nonintegers.

aweight provides analytical weights. Imagine that your data is collapsed including the mean of another variable, say wage. In that case the count still works as in fweight or pweight for point estimates, but precision increases with higher weights as the variance of the expected mean is more precise the more cases there are. Both pweight and aweight do rescaling not to inflate the number of cases above the total count, in contrast to fweight.

In sum, all weights return exactly the same coefficients, but different standard errors depending on the kind of data we're dealing with. One further note of caution: pweights and aweights are nonintegers, so precision is very important. I recommend storing such weights at double precision, not float. Also convert data from other formats using the 'double' option.

Links

dinsdag 3 februari 2015

Maximum Likelihood - with examples of regressions of continuous and discrete dependent variables

The idea of Maximum Likelihood is very simple, yet powerful. A regression returns a residual for every case, telling us how far the estimation is off. In Ordinary Least Squares, we minimize the residual. In Maximum Likelihood, we ... maximize the likelihood of the prediction. It sounds like it is the same thing, and in some cases it is the same thing. But whereas OLS mathematically computes the parameters of the model, ML uses a maximization algorithm (generally Newton-Raphson) that checks multiple possibilities. These are called iterations, and when the model works well, they will 'converge'. If there is no maximum (i.e. the maximizand is non concave), or several maxima, you are out of luck.

Continuous dependent variables

Take a simple regression like y = Xb, and suppose - as in simplified OLS - that the error term is normally distributed. Then e = y - Xb.

Now let f(e) = (2.pi.sigma^2)^-(1/2) . exp(- (e^2)/(2.sigma^2)

In this case f is a normal density function or normal distribution. If the error is 0, the density is highest (around .44), so a good estimation will need to have small errors.

The likelihood function we want to maximize is the joint probability distribution of the sample. This means we compute the density of the error for every case and multiply all of them. We may just as well take the log and sum the logs of the error density. When this sum of logs is maximizes, we have the same set of coefficients b as in OLS. Note that this will be the true b only if the error term is indeed normally distributed.

Discrete dependent variables

When the outcome variable y_obs is binary (0,1), such a linear model will not work. Instead we estimate a function that lies behind the observed outcome. If that function results in a number greater than 0, we will observe a success (1) and if not, a failure (0). Hence again, we have:

y = xb

Yet in this case we cannot simply compute a residual. In the probit case, we take the cumulative probability of xb. It is obvious that this depends on a scaling parameter sigma.

The likelihood function is then P(xb) if y_obs = 1, and 1-P(xb) if y_obs = 1. Another way to express this is:

L_i = P(xb)^y_obs * (1-P(xb))^(1-y_obs).

Maximizing the product of all L_i or the sum of the logs yields good estimates of b.

Note: an alternative to probit is the logit or logistic regression. Then the function P is not the normal cumulative distribution function, but instead it is the inverse logit:

P(xb) = exp(xb) / (1+ exp(xb) )

The advantage here is that no scaling sigma has to be estimated. It used to be popular in the early days of sociology, before econometrics became dominant in the social sciences.

Links

The last link has a nice picture comparing the cumulative distribution function and the inverse logit (or logistic) function. The latter has 'fatter' tails, even though the difference is negligible in practice.








zaterdag 27 december 2014

Vim

Vim or MacVim, a marvellous piece of software I have discovered too late but at least.

It is a text file editor. Do not expect anything else. It's like notepad, but with functionality. It highlights programming syntax like Notepad ++, but it is also on Mac.

You need to understand the two modes:
  • Edit
  • Insert
The Insert mode is your ordinary text editor. You can type in this mode. The Edit mode is everything else you do with the text, such as copying and pasting, finding and replacing strings, 

You activate Insert mode by pressing -a- and leave it by pressing -esc-. 

Here's the function(s) I have been using:

vrijdag 25 april 2014

Variance eaten up by extreme means (truncation)

If scales are continuous and limitless, variance and means are two different things. If scales are limited, however, this is not the case: towards the bounds, there will be substantially less variation than toward the center of the scale because of censoring.

The issue arose when we wanted to see whether there was divergence or convergence of job quality in Europe. If the scale was wide, the evolution of the variance would tell this, but if the scale is limited and the average moves towards one of the bounds, the variance will be wrongly considered to indicate convergence.

As a solution, I would compute some kind of one-tail variance on the longest tail if the other is strongly censored, hence assuming the latent distribution is symmetric. I have not seen such a measure yet, but it is easy to calculate.

In a way, it is a measure that should be possible to derive from truncated regression. It would be nice to do that.

Variance eaten up by extreme means (truncation)

If scales are continuous and limitless, variance and means are two different things. If scales are limited, however, this is not the case: towards the bounds, there will be substantially less variation than toward the center of the scale because of censoring.

The issue arose when we wanted to see whether there was divergence or convergence of job quality in Europe. If the scale was wide, the evolution of the variance would tell this, but if the scale is limited and the average moves towards one of the bounds, the variance will be wrongly considered to indicate convergence.

As a solution, I would compute some kind of one-tail variance on the longest tail if the other is strongly censored, hence assuming the latent distribution is symmetric. I have not seen such a measure yet, but it is easy to calculate.

In a way, it is a measure that should be possible to derive from truncated regression. It would be nice to do that.

dinsdag 29 januari 2013

Global Labour Column (ILO)

Short, insightful articles on work. Well worth reading!

http://column.global-labour-university.org/

woensdag 5 december 2012

Latin abbreviations

APA has provided this list, where you'll find that a.o.:
  • E.g. = exempli gratia = for example
  • I.e. = id est = that is
  • Cf. = confer = compare with
  • Etc. = et cetera = and so on
  • Well... et cetera.
Moreover, we should use al of these abbreviation only within parentheses (and footnotes maybe). Except for et al. (et alii), which is used all over the place. And mind you, when you use these abbreviations, there is no need to italicize!

Some more words of warning from the gentle folks at Sussex. Not.
And from Europe! Here.

vrijdag 21 september 2012

Labour Stats goes through a linguistic crisis

I thought about changing the name of this blog to 'Labor Stats'. Decent American spelling. It is said to be more logical. But just how logical do languages have to be? Does logic imply internal consistency, respecting etymology or rather streamlining common practice? Then which spelling is preferable: Merriam-Webster, Cambridge Dictionary, Oxford English Dictionary, BBC practice, MS Worst?

Having studied classical languages, I first considered etymology and the infamous -ising/-izing debate. The Greek ending is -idzein, Latin -izare. So why making it French when it isn't? Even old English uses the -izing form, although mixed with the s-spelling. There are other similar cases. Color is Latin too, it means colour. Why not simply write it the way it has always been (like the Americans do)? Merriam-Webster is on your side.

Second: internal consistency. Humour leads to humorous. We could spell it humor right away. It is more logical, easier for foreigners. In particular I like the use of the suffix -ize for everything you 'make'. That keeps you from writing analize, unless you have perverse motives. It is analyse or analyze. There's a preference for the former, because it comes from analysis-ize, largely omitting the suffix. Common British spelling (en-UK) thinks about it differently.

Etymology and consistency. Jolly good, but we may not forget that language is an independent cultural object. Even strict grammar and spelling rules cannot enter the living room of people and decided what dialect they speak, words they take up. During some period of time, French has influenced English. It has changed words English already had, and introduced words that have a history before they were French. These words had a history before they were Latin or Greek too. It is a cultural bias to assume civilization started 1000 BC. Maybe writing did, but then again this must have been more evolutionary than revolutionary. If you study classical languages, you quickly learn a new vocabulary for every Greek poet you read. Don't try to find the original writings, cause you will not even understand most signs.

This is exactly why I keep Labour in Labour Stats. It looks familiar, it has history and it's not that far from logical. If English would evolve towards 'labor', I would not mind at all, but for the time being labour ends on -our, error on -or and labourer on -er. It has been different before and it will be in the future, but the changes are minor and reflect the history of a word.

Then again, when the European Union tries to -ise English, primarily to please the French, I will first oppose. For me, the Oxford English Dictionary is a way more important authority that nicely balances logic, etymology and consistency. It interacts with language as it is used, without standardizing fads. Here is someone who agrees.

P.S. Don't try to make Microsoft Office talk Oxford English. You just bumped into the limits of closed source software.

dinsdag 11 oktober 2011

Alternatives to ... Powerpoint

I wasted some time looking around for Powerpoint alternatives. Basically, my conclusion is: don't do that. It's not because Powerpoint is such a great program, but rather because the alternatives have limitations which are more serious than the rightly criticized linearity of Powerpoint.

Anyway, what to expect:

  • [Sozi] It's open source and entirely vector based. It would make for a good banner or a graphic scheme you want to discuss focussing on the composing parts. *
  • [Vue] The best alternative in my view, which lets you construct different paths over the slides and insert midway overview in the presentation where you can get off road. A good idea, but not very user friendly. ***
  • [Prezi] This is very good looking software, but it has some major drawbacks. First of all, it's all flash and therefore heavyweight. Any animations you include need to be flash movies, which are very hard to make. The choice of slide templates is limited and customization options do not really help you out. There is a svg-alternative in the making, but that means it is currently useless. **
  • [ahead] As Prezi, but better looking. Spend some time and impress your audience. Not for quick tasks and only online. ****
  • [Zoho] As Powerpoint, but online and owned by Google. **