Rāda ziņas ar etiķeti groundwaters. Rādīt visas ziņas
Rāda ziņas ar etiķeti groundwaters. Rādīt visas ziņas

2011-11-29

long term monthly mean over 100 years

MS Excel is one of the most popular spreadsheet programs in the world. It is possible to write nice macros in VBA, it is handy and mostly meet a usual user needs. For me it is very convenient data storage type - but when it comes to work with large data I prefer writing macros or using external modifying. And here comes Python, again. Ok, the problem:
I have an excel document with data in one sheet, ~19 columns and daily measurements for 100 year period ~34000 rows. The A column contains date, other 19 columns counting from B - values/measurements. And I need to calculate firstly the long term monthly mean normal values for each column (resulting 12 values for each month) and secondly - percentiles over these 19 columns.
The reading from excel is done using XLRD package and percentile calculating - using scipy stats . These two libraries should be installed additionally to run my class.
The created class is so-so complete with a lot of space for possible improvements; currently it met my needs and I`m ok with that.
The feeling of satisfaction after doing such work manually in excel or creating script is completely different, and in the latter case it`s more positive.
This class is free to use, but if you are using this class, please, make the reference to me and let me know about the project where it has been used.

2011-09-17

Linear interpolation. Extended version.

I already wrote about the linear interpolation and related problems which were successfully solved. However, full convenience requires some monthly statistical capacities. Usually the data series, including interpolated ones, does not start and end with first and last day of the month, respectively and thus the first thing is to cut off the first and last incomplete month measurements. And only then it is possible to split the data series by month. Splitting occures in the loop splitting the corresponding date and getting the month and comparing it with the previously splitted month value. If the measurement is in the same month, the new two dimensional array is created with the date (year and month) in the first dimension, and (measured including interpolated) value in the second. And if the current month value differs from the detected month value in previous cycle, it is time to change the first - date dimension value to the next one. After creating such array, it is time to make some statistics for each of inner measurement array.
You can try the linear interpolator here and see the source code (not yet OOP, but i hope soon) - here.

2011-08-30

Linear interpolator script

Mostly measured data in databases are defined in exact time or date. But sometimes it is necessary to interpolate between different dates because of skipped measurements. For example, in earlier times, groundwater level measurements were made "by hand" after every three to seven days, depending of site and monitoring type. But nowadays groundwater has been monitored using different loggers, which are called - divers. These little gadgets gives possibility to get daily data.


The problem is that working with daily measurements and not taking care about skipped days, statistical weights could be different from the case, when in these skipped days the measurements are  interpolated.

I made a little script, where measurements and their corresponding dates are taken from mySQL database using PHP; and in the database there could be only measured data - this script interpolates the missing values.

2011-02-25

HeidiSQL

Very nice program to access remote SQL server.
Also, I`ve learned some neat things about MySql functions and extended use.
For example, here is some code
What does it do?
It`s looking for all similar months and returns an average value from all values on given month.
In the inner "select" I`m selecting all columns and I`m adding additional column, where the each month (meneshi) is added in each row using "where"
And then in the outer query I`m gruping values by newly created month column and calling an average function. It returns all average values for each group - month.
And lastly I`m ordering results by month. Easy as pie, actually.

Only thing I need to change is "ns" in third line from end of the script.
Next challenge would be to make average results for all ns.
select avg(lim), datums, meneshi
from
( select ns, lim, datums,
CASE
WHEN datums like '19__-01-__%' THEN '01'
WHEN datums like '19__-02-__%' THEN '02'
WHEN datums like '19__-03-__%' THEN '03'
WHEN datums like '19__-04-__%' THEN '04'
WHEN datums like '19__-05-__%' THEN '05'
WHEN datums like '19__-06-__%' THEN '06'
WHEN datums like '19__-07-__%' THEN '07'
WHEN datums like '19__-08-__%' THEN '08'
WHEN datums like '19__-09-__%' THEN '09'
WHEN datums like '19__-10-__%' THEN '10'
WHEN datums like '19__-11-__%' THEN '11'
WHEN datums like '19__-12-__%' THEN '12' END as meneshi
from hmetlimeni where ns=9664 ) as tble
group by meneshi
order by meneshi

2010-03-04

long time no see.

I`m involved in some new projects dealing with groundwaters, GIS and modelling. I`ll try to give some updates here how the things are going, problems and possible solutions.
Groundwater. Freeware. Modelling. Nice pictures.
These are keywords or my work. I started with all of them together reading materials and finding out what kinda free software I can use.
The OS and software I`m on right now -
OS = Linux Ubuntu 9.10 (Karmic Koala).
Software. I`ll focus mainly on QuantumGIS, GRASS and R, following discussions with collegues in Uni. But for each of them there will be a seperate story and discussions.
Here we go.