Showing posts with label Ideas. Show all posts
Showing posts with label Ideas. Show all posts

Thursday, March 01, 2007

Significant paper

@article { Lee2005,
title = "Real observations of coastal algal blooms by an early warning system",
author = "J.H.W. Lee and I.J. Hodgkiss and I.H.Y. Lam",
year = "2005",
journal = "Estuarine, Coastal and Shelf Science",
volume = "65",
pages = "172--190"
}

paper regarding the hongkong evaborate study.

Wednesday, February 14, 2007

The type of inference

In any kind of choice between techniques, it is important to know the type of inference we want to make. There is no universal solution, because there is no loss-less generalised answer.

Thursday, January 11, 2007

How much is too much in sampling? How much is enough?

Thursday, December 14, 2006

predicting change or predicting absolute values



we define
forecasting as testing model on data not utilised to develop model and
predicting as testing model on data which is obtained from observing the system in a future time.

in that case, the above graph shows a prediction of algal biomass one hour ahead of time. the system was developed to emulate the natural function relating (water quality parameters) and (chlorophyll one hour ahead in time) as observed in a period of 388 hrs, which is around 16 days and 4 hrs. it is validated over the next (approximately) 16 days and is tested over the next 32 days. The figure below summarises this.

what we find is that we follow trends well, but base value is lost. which suggests that we might as well try to predict 'change' in algal biomass. we could experiment by defining change as a vector value - with a magnitude and a direction.

One reason why this idea has not been persued yet is also because the above graph collapses to almost gibberish when time gap is increased further - see below.

To be fair, the above graph is not ALWAYS the case, the exact graph changes often enough - in one case it was tracing ok for a while and then turned into a straight line. but in all cases the regression coeff drops to something like .3 and the graph whatever be their flaw they all have this in common that they ARE NOT ACCURATE!

Thursday, November 30, 2006

Comparing Models - 2

The question is - is the error being amplified or is the accuracy being amplified?

in high variance systems, it appears, that the model ends up emulating different sections of the data set. From the point of veiw of the statistical measures of accuracy, models with very different ____ qualities, may appear equivalent.

Since this situation would always be reflected in higher error in atleast one of the three error values, we at least know when the model is definitely incomplete. an objective measure of completeness is not easily found because we do not have information other than the training data (which is called - lack of meta data) to compare it with. an issue resulting from dealing with a largly unknown system.

Regarding Sensitivity Analysis
if similarities are found between complete models and incomplete models -
can it be concluded -
that the similarities are strongly persistent in the entire set.

(data and random nos. should not give the same kinds of results - ... does this need any more work to be done.

eventually, the results of sensitivity analysis is dependent on
- the raw data,
- the neural network model

Wednesday, November 29, 2006

Comparing models

The difference between the converged networks (or models where values for free parameters have been determined) may be more for certain datasets (DS4); and less for others (DS1).
what i mean is,
for DS1 - after repeating the process a reasonable no. of times, the network tends to reach some kind of a minima region where networks are very similar - this region seems to have close approximations of the actual relationship.

on other hand for DS4 - say, two networks converge and give very similar regression coefficient values (and sometimes even similar mse and mape values) but there is a huge difference between these networks. what is interesting is that they give very similar sensitivity analysis results.

Updates:

for example of what i am saying
- see ds4fss2days3hrsHL7_no2 and ds4fss2days3hrsHL7_no3

what, i guess, i am questioning is how well do 3 values describe the quality of the model. do they do it well - because that would mean that two models that visually look very different from each other will be the same quality. and how is it affected by consistent results from SA or the lack of consistency in results from SA.

Monday, November 20, 2006

Objectives of data analysis

Why analyse data?

you try to answer the question - is there a hidden determinism in your data?
and after knowing that you would like to
a) predict or
b) extract a deterministic signal from noisy background
c) gain better insight and understanding of the underlying dynamics

paraphrased from Chaos and Time-series analysis by Sprott.
(a similar thing is mentioned in kingston - see.)

My ideas -
Step 1:
Is there a hidden determinism in the data?
Traditional statistical technique -
Autocorrelation?

Itertative NN technique -
Test existence of the relationship using alternating division in data and using a large test set.

Step 2:
Gain better understanding of underlying dynamics:
is there a periodicity?
is there relationships between parameters?

Traditional statistical technique -
Periodicity -
Fourier analysis
Lyuponov exponents?

Relationships between parameters –
descriptive statistics – scatter diagram
ANOVA / MANOVA/

Iterative NN technique -
Periodicity –
??

Relationship between parameters -
Non – linear principal component analysis?

Sensitivity analysis
weights method shows that parameters are highly dependent
derivatives method shows that chlorophyll is more sensitive to changes in certain parameters. (get exact statement)

Step 3:
Is there a predictive function/model/law?
Traditional statistical technique -
MA?
ARMA?
ARIMA?

Iterative NN technique -
Test existence of a predictive function using sequential division in data (and using a large test set?)


Before we get into any further tradition statistical technique selection – check assumptions.

a) what kind of variables – ordinal, continuous etc are required and
b) what kind of distribution is required – normal? whatever..
c) how many independent and dependent variables are accounted for.

Sunday, November 19, 2006

On matching techniques and problems

In Kanal (1993), the following figure is shown to present the various techniques used for various aspects of pattern recognition. It is suggested that we may look at the sceanario as a "bag of tools for a bag of problems".


My point is we really need to see what is the limiting factor here - if data is the limiting factor then using fancier technique would not help - and therefore more than one technique is sufficient.*
If we do not observed an entire cycle of the process then we cannot expect the fancier techniques to help. The problem really does boil down to knowing if we are observing at the right temporal scale.



Another diagram in the paper which is of relevance (and is presented wrt a case study)


*it is difficult to see how can one argue against the other techniques if they have not even been applied (specially since all these papers argue that each technique and each problem need to be matched; there are no general solutions for all complex problems). however, i wonder what would be use the use of the above arguement if all techniques are applied to test its validity.

Monday, November 13, 2006

What can I infer from 'Results of Statistical Tests'?

Statistical results done on data throws up many patterns much like the data itself. It is getting interesting as I try and figure out

- which of the results are showing a pattern because of the pattern inherent in the statistical test. In my case, especially in SA by partial derivation method

- which are being shown because of extreme values present in the data set. In my case, especially when the point keeps moving between training, validating and testing data sets.

(and the most brilliant one)
- how much of it getting stuffed up because I am using the wrong scale to look the environment.

Sunday, October 15, 2006

On the position of ANN models in population ecology

Black boxes: If models that are not designed to reflect the actual processes in the under-lying systems are called black boxes. Then, in that respect ANN are indeed black boxes. ??check

Models or Laws: Model is a relation or a set of relations that relate the various components of a system. There are two ways to reach a model
- one is to study the physical and chemical processes and then use mathematical statements to express these processes. The mathematical statements would be called models; and
- two is by studying various components and finding the statistical relationships between them. The statistical relationships would be called models.

could both the methods be called empirical?

in any case, from kingsland's book pg. 100 Thompson's hypothesis is presented
If ecological interaction were found to display an underlying regularity, and if this regularity could be described mathematically, then mathematics might serve as a theoretical basis for population ecology.


Also, the book mentions on pg 85 that the term laws (and may i add models, as well) has been used in more than one way -
Specifically, Pearl used the term 'the law of population growth' for the logistic curve to possess universal applicability. Where as, Lotka used the same term
to mean an empirical relation between events having no apparent connection to principles of a more general nature. By this criteria, any other equation fitting the observations would be qualified equally as a law.

Lotka perceived that an empirical law of this sort imposed limits in two ways. First, because the fundamental principles underlying the curve were unknown (MARK1), the exact form of the equation had to be determined anew for each examples. Second, it was not possible to extrapolate much beyond the observed events, because unknown factors might come into play outside the observed range and cause departure from the law

Talking about the logistic curve as the law of population growth - the value lay not in universality or predicting but
in the fact that it could be so easily derived from first principles, and that its constants r and K were biologically meaningful. The equation represented in other words an argument about the population which could be useful as tool of research.- MARK2

Lotka finds use of such a law in this way
"An empirical formula is therefore not so much the solution of a problem as the challenge to such solution. It is a point of interrogation, an animated question mark."
This is later (Pg 87) explained as -
by looking at how a population departs from the law, one may get a more realistic idea of the actual mechanism underlying the population growth... Knowing how a population deviates from the law tells one how to refine the initial assumptions to get more accurate understanding of how populations behave.


By reading MARK1 and MARK2 above, what I am trying to do is see if there is a fundamental similarity between the logistic curve and ANN model as population law. If so, then we know exactly where to place the ANN models in the study called population ecology or population dynamics of phytoplankton [Note to self - there needs to be a discussion on exactly which term would I be using]. And, thus how far can the results be extrapolated.

My current hypothesis (??) is -
By developing a series of empirical models and using those empirical models to
- evaluate the technique (the model is objectively evaluated by having a large percentage of testing data)
- study how the relationship varies (by performing Sensitivity analysis on model that have been found to be good approximation using objective measures)

the next things -
1) look at what exactly is Thompson above is talking about.
2) See exactly what assumption and limitation are associated with the SA techniques that I have used.
3) A very brief discussion on which term (phytoplankton, blue-green algae, any other) would be used and why.

Friday, October 13, 2006

Research Interests

Since I am looking for a job, I would like to work, or rather continue working but at a larger scale in the area that I am already working in, Population ecology.

I am looking at the population dynamics of phytoplankton in an estuary in Victoria, using bio-physical time series. I have tried to understand the short time scale dynamics, varying from the order of hours to days. I have used an iterative non-linear statistical modelling tool which in other words is called feed-forward neural network. (neural networks have been given the bad name of being a black box, however I think it has got more to do with they way they are used).

Using partial-derivative sensitivity analysis (I have come across at least one paper which gets upset with this terminology) I have tried to evaluate the strength of the relation between each of the bio-physical parameter and chlorophyll. (chlorophyll is used an indicator of phytoplankton population.) And to see how the relationships change as the time interval between chlorophyll and bio-physical parameters is increased.

Now what I would like to do further is

  1. Study non-linear time-series and see how does chaos (which appears in weather time series and water flows) appear in ecological time-series.
  2. Study oceanography, my approach to my project was as a mathematician. I have, hopefully, learned some decent amount of biology in the process. However, I am sure there are quite a few things out there – so, I would like to do a short course in oceanography. (Note: I did NOT say I wanted to do a short course in chaotic time series analysis)
  3. Study Bayesian: Bayesian is used to quantify uncertainty. It is being used to make ecological models more accessible to decision makers who can use them as decision support systems.