Posts tonen met het label Python. Alle posts tonen
Posts tonen met het label Python. Alle posts tonen

woensdag 23 september 2015

Programming languages compared, and why I'm sticking with Python until Julia grows up

Comparisons of programming languages abound, especially with regard to running speed. This paper in the Journal of Economic Dynamics and Control by Boragan Aruoba of the University of Maryland and Jesús Fernández-Villaverde of the University of Pennsylvania is yet another one, albeit in a peer-reviewed journal and with a procedure (value function iteration) that is common in my field. Main observations:
  • When it comes to speed nothing beats C++ and FORTRAN.
  • Julia performs really well: only 2.37 times the running time of C++.
  • Python and R are slowpokes, at about 45 (Python with the Pypy interpreter) to 491 times (R without its Compiler package) times the running time of C++.
  • Matlab is somewhat inbetween the slowpokes and the frontrunners, with about 9 times the running time of C++.
  • R, Matlab and Python can get a boost from just-in-time-compilers and C compilers (like Rcpp for R; or Numba or Cython for Python) that make their running time comparable to that of Julia (although Rcpp is still a bit disappointing: 5.4 times the C++ running time).
Adding my own considerations:
  • I tried C++ and gave up very quickly. I bloody hate it with a passion.
  • I find R not much better: even though it is a scripted language it is clumsy, illogical, and makes for horrible code. Rcpp worked fine until I found out that it cannot handle matrices with more than 2 dimensions. Speeding up your R code with bigger matrices requires the use of full-blown C++. In other words, two languages that I bloody hate.
  • I value my independence, so I prefer to work on my own laptop with my own licenses. That makes Matlab, with its steep license fee, a no-go.
  • I've experimented with Cython and it seems to work quite well. I love the accessibility and clear layout of Python; moreover, my university teaches Python twice a year in an undergrad course. The only real problem with Python is that it is reputedly difficult to parallellize.
  • One day I'll start using Julia. It's fast, accessible, and (so they say) easy to parallellize. But not until there is a stable version and a decent IDE.

maandag 24 november 2014

Value Function Iteration in R, Python, GAMS

So I've spent a couple of evenings debugging my code, trying it in different languages to see whether it was me or the built-in functions I used (it was me), so I ended up with the same model in R, Python, and GAMS. I figured I might as well post the code on the web.

Yes, I found the problem, and no, I'm not going to tell what it was because I find it too embarrassing.

zondag 27 juli 2014

Judd and Guu's stochastic perturbation model in Python

I just uploaded a Python version of Judd and Guu's perturbation code to my website (code; notes).

If you happen to be a Python programmer, or a computational economist (even better: both!), then any feedback you can give on the code is highly appreciated. I'm fairly new to Python, so I'm sure I could have programmed some parts more efficiently.

Oh, and here is a picture of a bunch of Canberran kangaroos:


I thought you'd like to know.

woensdag 23 juli 2014

In case you were wondering what I'm doing in the Australian winter

I'm here for two very interrelated goals.

I'm having another assessment meeting in September 2014; this time it's about a possible promotion to associate professor. So Goal #1 is to take a good look again at my research and education vision, and discuss it with whoever I can discuss it with. I got quite some inspiration from the keynotes and discussions at IIFET2014. Not that I went there with a blank slate, but it was good to see my ideas confirmed, in a way, and complemented by other people's ideas.

I have decided long ago that I will focus on the economics of coastal and marine ecosystems. My background is mainly in bioeconomic modelling, so it is logical to focus my research on the kind of questions that require such modelling. But then the question arises: aren't many other people doing that already? People have been doing theoretical fisheries economics since the 1950s (or longer, if you consider Jens Warming's work). And there are gigabytes of applied bioeconomic fisheries models like FishRent and Mefisto, and wholesale ecosystem models like Atlantis, where fishers are included as some sort of predators.

But that's it, actually: either the models are very abstract and qualitative, so that they can be analysed on paper, or they are very detailed and quantitative, so that they can be used for policy assessment or scenario analysis. The problem with the first is that they lack realism; the problem with the second is that they lack transparency. Either you can explain what drives your results, but then your results are close to useless for policy makers, or you can advise policy makers but you cannot explain where your advice comes from.

What has not yet happened much (I know there are people doing it, but not many), is to take the theoretical models, and make them more realistic to the point where you can maintain some intuition as to what drives your results, even though you cannot prove fancy theorems anymore. Macroeconomists and financial economists have reached that stage long ago: where their models get too complicated to be solved by some math magic, they use computation. This way you can add more realism, while maintaining a fair amount of insight into the mechanisms at work. My intention is to apply such computational methods to problems with coastal and marine ecosystems. This includes a lot of fisheries, but also other ecosystem uses, goods, and services.

Which brings me to Goal #2. The Crawford School of Public Policy of Australian National University has among its staff a number of people who have applied computational economic tools to fisheries problems, like Tom Kompas, Hoang Long Chu, and Quentin Grafton. I'm here to learn at least some of the methods they use. Originally I wanted to stay about two months, but for several reasons I only have about two weeks. But in the short time frame I have I'm trying to get the most out of it.

And lo and behold, I have a first result to show you. My first hurdle was writing a perturbation model in a program I can work with. Their models run in a combination of Matlab and Maple, but I don't have a license for either of them, and I'm not well-versed in Maple. Hoang Long Chu was so kind to give me a paper by Kenneth Judd and Sy-Ming Guu on writing perturbation models in Mathematica - another program I don't use, but luckily the paper explains the method well and it presents the entire Mathematica code for a simple optimal growth model. So I decided to write the same method in Python - my language of choice for its elegance, simplicity, and speed (ok, compared with R, which is neither elegant, nor simple, nor speedy). It took me a few days but here it is: the Python code and a pdf with some notes on the paper and the model.

donderdag 3 april 2014

A simple Python script to make a literature table

Geeky post again - no math this time, but computer code.

I'm sure people have done this before, but I thought it would be a nice opportunity to practice my Python skills to write a small script for the following problem.. Usually when I read a scientific article I watch out for the following elements:
  • Innovation: what does the study do what others haven't done before?
  • Method: what method did they use?
  • Data: where did they get their data from?
  • Results: what are the main results?
  • Relevance: who benefits from this research, and how?
I also like to place the research in one of the four quadrants in this post. I find it helpful to make an overview of these questions in a table:

1st authorYearJournalQuadrantInnovationMethodDataResultsRelevance
Kompas2005Journal of Productivity Analysis4Estimates efficiency gains quota trade for Southeast Trawl Fishery, AUStochastic frontier analysisAFMA and ABARE survey dataITQs gave efficiency gainsPolicy debate on ITQs
Kompas2006Pacific Economic Bulletin3Estimates optimal effort levels and allocation across speciesMultifleet, multispecies, multiregion bioeconomic modelSPC dataEffort reduction needed; optimal stocks larger than BMSYPolicy debate on MEY

But here's the problem: I usually make my notes in a bibtex file (as a good geek should), which looks like this:

@ARTICLE{Kompas2006PacEconBull,
  author = {Kompas, T. and Che, T.N.},
  title = {Economic profit and optimal effort in the Western and Central Pacific tuna fisheries},
  journal = {Pacific Economic Bulletin},
  year = {2006},
  volume = {21},
  pages = {46-62},
  number = {3},
  data = {SPC data},
  innovation = {Estimates optimal effort levels and allocation across species},
  quadrant = {3},
  keywords = {tuna; bioeconomic model; optimisation; Pacific},
  method = {Multifleet, multispecies, multiregion bioeconomic model},
  results = {Effort reduction needed; optimal stocks larger than BMSY},
  relevance = {Policy debate on MEY}
}

@ARTICLE{Kompas2005JProdAnalysis,
  author = {Kompas, Tom and Che, Tuong Nhu},
  title = {Efficiency gains and cost reductions from individual transferable quotas: A stochastic cost frontier for the Australian South East fishery},
  journal = {Journal of Productivity Analysis},
  year = {2005},
  volume = {23},
  pages = {285-307},
  number = {3},
  quadrant = {3},
  data = {AFMA and ABARE survey data},
  innovation = {Estimates efficiency gains quota trade for Southeast Trawl Fishery, AU},
  keywords = {individual transferable quotas; stochastic cost frontier; fishery efficiency; ITQs},
  method = {Stochastic frontier analysis},
  relevance = {Policy debate on ITQs.},
  results = {ITQs gave efficiency gains}
}
I don't want to copy it all by hand, so I wrote this little script in Python to convert all entries in the bibtex file to a csv file:

import csv
from bibtexparser.bparser import BibTexParser
from dicttoxml import dicttoxml
from operator import itemgetter

def readFirstAuthor(inpList,num):
    author1 = ""
    x = inpList[num]['author']
    for j in x:
        if j != ',':
            author1+=j
        else:
            break
    return author1

def selectDict(inpList,name):
    outObj = []
    for i in range(len(inpList)):
        if name in inpList[i]['author'] and \
            inpList[i]['type']=='article':
            outObj.append(inpList[i])
    return(outObj)

def selectFieldsDict(inpList,fieldNames):
    outObj = []
    for i in range(len(inpList)):
        temp = {}
        for n in fieldNames:
            if n == 'author':
                author1 = readFirstAuthor(inpList,i)
                temp['author'] = author1
            else:
                if n in inpList[i]:
                    temp[n] = inpList[i][n]
                else:
                    temp[n] = 'blank'
        outObj.append(temp)
    return(outObj)

fieldnames = ['author','year','journal','quadrant',\
    'innovation','method','data','results','relevance']

with open('BibTexFile.bib', 'r') as bibfile:
    bp = BibTexParser(bibfile)
    
record_list = bp.get_entry_list()
record_dict = bp.get_entry_dict()

dictSelection = selectDict(record_list,'Kompas')

fieldSelection = selectFieldsDict(dictSelection,fieldnames)


test = sorted(fieldSelection, key=itemgetter('year'))


test_file = open('output.csv','wb')
csvwriter = csv.DictWriter(test_file, delimiter=',',\
    fieldnames=fieldnames)
csvwriter.writerow(dict((fn,fn) for fn in fieldnames))
for row in test:
     csvwriter.writerow(row)
test_file.close()

If you are a Python developer: any comments on this are welcome. I'm sure it's not perfect.

zondag 6 januari 2013

Why (for the time being) I'm sticking with R

I'm a big fan of open source software. OK, I know the Dutch have a reputation for being stingy but let's face it: much of the software we use in economics (Stata, Matlab, Maple) is terribly expensive. So the only time I can use these programs is at the office (which, I admit, should be considered a healthy thing). To be able to work on my laptop when I'm at home (or in a hotel room, or in an airplane, for that matter) I try to work as much as I can with their open source equivalents as much as I can.

One of the programmes I've been using is R (a horrible name to Google for by the way), but in a sort of on-and-off way. It is less user-friendly than Matlab, much slower than Matlab, and contains fewer possibilities for statistical analysis than Stata. So I'm still fiddling around with programming languages like C++ (probably even faster than Matlab, but rabidly user-hostile) and Python (more user-friendly than C++, and perhaps as fast as Matlab) for calculations.

Slowly, however, I'm coming round to R, in my teaching as well as in my research, for a number of reasons:

  • Marine biologists use it a lot, and using the same software helps the communication - it also makes it more likely that you can ask a close colleague how this @#%! package works.
  • By the same token: some of my students, i.e. those who have taken marine ecology modelling courses, know it already.
  • I can use it in my environmental valuation classes (statistics) as well as in my resource economics classes (modelling), so that again, some students in one course know it from another course I'm teaching.
  • It seems that R finally has a decent package to do conditional logit and probit (or, as others call it, alternative-specific multinomial logit and probit).

If only they could make it a lot faster, because it is too slow for value function iteration.