Sunday, January 16, 2011

Too frequent update of Symantec Endpoint Protection in Mac

I have installed Symantec Endpoint Protection (SEP) for Mac in an unmanaged mode. The problem is its LiveUpdate runs too frequently and I can't disable it or modify. Although it provides Symantec Scheduler UI, it doesn't look like working at the beginning.

It turns out SEP installs a default schedule owned by a root user, which a normal user can't see in his/her Symantec Scheduler UI. A solution is as follows:
1. Open a terminal
2. Type sudo symsched -d all
3. Setup a new schedule by using Symantec Scheduler UI or a command like,
symsched LiveUpdate "Update All Daily" 1 1 -daily 13:00 "All Products" -quiet

Reference:
http://www.symantec.com/connect/forums/sep-mac-live-update-bouncing-dock
http://www.symantec.com/business/support/index?page=content&id=TECH134203&locale=en_US
http://www.symantec.com/business/support/index?page=content&id=TECH105502

Sunday, January 09, 2011

Time measurement methods

I've just recently tried to decide what is the best way to measure times. The followings are a short summary I got after browsing a few sources in the Internet (See references).

1. gettimeofday
-. Resolution : 1 microsecond
-. Have a negative time issue due to clock reset or NTP

2. clock_gettime(CLOCK_MONOTONIC)
-. Resolution : 1 nanosecond
-. Use HPET, an external chipset on board.
-. An optimal choice, considering both reliability and portability

3. RDTSC
-. Resolution : 1 nanosecond
-. Use RDTSC, a cpu instruction
-. Very light-weight but lots of issues on reliability
-. For a short run and if it is clear on the reliability issues (see Ref 2.), it's good to use.

References:
1. http://aufather.wordpress.com/2010/09/08/high-performance-time-measuremen-in-linux/
2. http://juliusdavies.ca/posix_clocks/clock_realtime_linux_faq.html
3. http://www.kerrywong.com/2009/05/28/timing-methods-in-c-under-linux/

Wednesday, July 07, 2010

HDF5 MPIIO's Hyperslab_by_chunk.c error in Mac

In Mac, I got a runtime error after compiling hyperslab's chunk example, Hyperslab_by_chunk.c

As mentioned in here, the error is Mac-specific but the solution what I found is a little different.

Use
hdf5_cv_mpi_complex_derived_datatype_works='no',
instead of
hdf5_mpi_complex_derived_datatype_works='no'

I have tested with HDF5 version 1.8.5.

Monday, November 02, 2009

A few useful Windows HPC commands

0. Set Default scheduler
set CCP_SCHEDULER=headnode_name

1. clusrun
One can run the following command from his desktop to setup compute node environments

clusrun /user:domain\username dir
clusrun /user:domain\username xcopy \\headnode_name\dir .
clusrun /user:domain\username setx PATH "somepath"

2. job
job list /scheduler:headnode_name

3. node
node list /scheduler:headnode_name

5. cluscfg
cluscfg setcreds


Sunday, November 01, 2009

Imporve serialize() in R for windows

serialize() function of R is very slow in Windows. My observation is that calling realloc is the bottle neck. The workaround can be replacing realloc with Rm_realloc in $R_HOME/src/main/serialize.c. More specifically, I have updated as follows:

1. In function, resize_buffer(...), replace
mb->buf = realloc(mb->buf, newsize);
with
mb->buf = Rm_realloc(mb->buf, newsize);

2. In function, free_mem_buffer(...), replace
free(buf);
with
Rm_free(buf);


The following is a short result running on my Windows 7 desktop:
The original:
> system.time(serialize(matrix(0, 1000, 1000), NULL))
user system elapsed
5.74 4.39 10.15
> system.time(serialize(matrix(0, 2000, 2000), NULL))
user system elapsed
85.40 74.80 161.62

After updating:
> system.time(serialize(matrix(0, 1000, 1000), NULL))
user system elapsed
0.78 0.30 1.10
> system.time(serialize(matrix(0, 2000, 2000), NULL))
user system elapsed
6.21 4.13 10.54

Saturday, October 31, 2009

Compile Rmpi with Windows MPI (HPC pack)

Compile Rmpi (http://www.stats.uwo.ca/faculty/yu/Rmpi/) in Windows

1. Install HPC Pack 2008 SDK with SP1 and modify line 314 in mpi.h (installed_dir\Include) as follows:
//typedef __int64 MPI_Offset;
typedef long long MPI_Offset;

2. Download Rmpi from http://www.stats.uwo.ca/faculty/yu/Rmpi/download/linux/Rmpi_0.5-7.tar.gz

3. Untar the source and modify src/Makevars.win by changing directory and library option (-lmsmpi) as follows:
PKG_CFLAGS = -I"C:\Program Files\Microsoft HPC Pack 2008 SDK\Include" -DMPI2 -DWin32
PKG_LIBS = -L"C:\Program Files\Microsoft HPC Pack 2008 SDK\Lib\i386" -lmsmpi

4. Compile and install
R CMD INSTALL Rmpi

5. Build for re-distribution
R CMD build --binary Rmpi

Wednesday, October 28, 2009

Setting up a non-admin SVN repository shared with multiple users

A trick is using SSH with public key as described in

As a quick summary, I have set up a SVN repository as follows:
1. create a svn root directory, which will share with others, and create a repository:
$ mkdir /path/svnshare
$ svnadmin create /path/svnshare/project

2. Get public keys of user to share svn repository and insert to ~/.ssh/authorized_keys the following line:
command="/path/to/svnserve -t -r /path/svnshare --tunnel-user=[USERID]",no-port-forwarding,no-agent-forwarding,no-X11-forwarding,no-pty [KEY_TYPE] [PUB_KEY] [COMMENT]
[.] should be replaced with users' information

3. Now users can access the shared svn repository remotely as follows:
svn co svn+ssh://[MYID]@[SERVER IP or DNS]/project

Note that use relative path names of repository after server ip or dns

Monday, June 29, 2009

More tips for using R with GotoBLAS

After building R with GotoBLAS (See http://jychoi-report-cgl.blogspot.com/2009/04/compile-r-with-gotoblas.html), a few things we can do for verification.

1. Download a R benchmark script (http://r.research.att.com/benchmarks/R-benchmark-25.R) and run it to check if every step works ok. (You can compare the performance with ones from normal R build too.)

2. If R is hang in calling eigen() function, try to rebuild GotoBLAS and R by using the same fortran compiler.

3. If nothing works, one can use ATLAS instead of GotoBLAS

Friday, April 24, 2009

Compile R with GotoBLAS

I want to share my experience to use GotoBLAS as an external multithread library of R. I didn’t make through performance test with other libraries (such as ATLAS) but I did got lot of performance gains with GotoBLAS in using R.

1. GotoBLAS from http://www.tacc.utexas.edu/resources/software
Follow instructions in 02QuickInstall.txt.

2. CBLAS from http://www.netlib.org/blas/blast-forum/cblas.tgz
This is not required for using R but you may need this for using GSL(GNU Scientific Library)
a. If the architecture is Linux, type
$ ln –s Make.LINUX Make.in
b. In Make.in, Modify BLLIB, CBDIR and –fPIC –lpthread to LOADER option.
c. Type make all for building libraries and testing
d. After completing, go to lib/LINUX and type
$ ld -melf_x86_64 -shared -soname libgotocblas.so -o libgotocblas.so cblas_LINUX.a

Note. try to use –m64 or –m32 if you are working with powerpc

3. R from http://cran.r-project.org/
Run configure as follows:
$ export GOTOBLAS_LIB=PATH/TO/GOTOBLAS_LIB
$ mkdir build_goto; cd build_goto;
$ ../configure --prefix=$HOME/usr/R/ --with-blas="-L$GOTOBLAS_LIB -lgotoblas -lpthread" --enable-R-shlib --enable-R-static-lib --enable-BLAS-shlib

Tuesday, April 07, 2009

Create AVI from R plot

1. In R, save plots as postscript (or png) files by adding sequence number. For example,

postscript(sprintf("%04d.eps", i))

2. By using ImageMagick’s convert command line tool, convert images from eps to png (You may skip if you have png files)

for f in *.eps; do convert -rotate 90 $f png32:$f.png; done

You may want to add the following options:

-resize 1280x720 : change image size
-bordercolor white -border 0x0 : add white background

3. Create AVI by using ffmpeg

ffmpeg -r 15 -sameq -i %04d.eps.png out.avi

You can control frame rate (-r rate) and quality (-sameq or –b bitrate)

 

Note:

If you need speed-up in conversion, you may try to use a bash script parallel.sh in http://pebblesinthesand.wordpress.com/2008/05/22/a-srcipt-for-running-processes-in-parallel-in-bash/

Sunday, February 01, 2009

Summary of the recommendation system survey paper

I’ve found a good survey paper about recommendation systems as follows:

Gediminas Adomavicius, Alexander Tuzhilin, "Toward the Next Generation of Recommender Systems: A Survey of the State-of-the-Art and Possible Extensions," IEEE Transactions on Knowledge and Data Engineering, vol. 17, no. 6, pp. 734-749, June, 2005.

A short summary:

-. A short definition of recommendation system: a problem of extrapolation for predicting unknown values

-. 3 approaches: i) Content-based, ii) Collaborative, and iii) Hybrid

i) Content-based RS(Recommendation System): a user’s feature is computed solely based on the user’s activity history. Can have the following limitations:
a. Feature extraction can be hard in some domain, such as multimedia or image
b. Over specialization: Diversity is required. Randomness, genetic algorithms, or some adjustment (remove too similar, or too different outputs)
c. New user problem: No information to consider

Known algorithms: (Naive) Bayesian classifier, Rocchio, winnow, ANN, …

ii) Collaborative RS: a user’s feature is computed by a group of like-mined people or peers. Limitations:
a. New user problem [83][89] : The same with content-based RS
b. New item problem
c. Sparsity: a few workaround ideas -- use of demographic information, dimension reduction, …

Known algorithms: clustering, Bayesian network, SVD, maximum entropy, …

iii) Hybrid RS: utilize both content-based and collaborative system.

Sunday, December 21, 2008

Flash 3D Engine: Sandy3D Vs. Papervision3D

Inspired by an article compared performance between Away3D vs. Papervision3D, I’ve just wanted to compare simple performance Sandy 3D (3.1 AS3) vs. Papervision 3D (2.0. Revision 849. Code name: Greate White) in rendering 1,000 objects.

As a result, the papervision3D is faster than Sandy3D by roughly 2~3 times or even more. My benchmark implementations of both Sandy and Papervision3D and source codes are available. 

Optimization advice of Sandy can be found here.

How to Parallelize

I’ve found a very good introduction about how to parallelize applications from http://www.cs.princeton.edu/courses/archive/spr08/cos598A/parallelization_course.pdf

Especially, a few tools are very helpful to analyze the code at the beginning stage:

1. gprof – profiler. Help to decide which part should look at. Consider to use a pthread wrapper available at http://sam.zoy.org/writings/programming/gprof.html

2. helgrind – To detect race conditions

Tuesday, December 09, 2008

GCC 4.3.2 Compilation and OpenMP

 

In order to try OpenMP3.0, I had been trying to install gcc-4.3.2 on an x86_64 Linux box without any luck. I got the following error in building gcc-4.3.2

/usr/bin/ld: crti.o: No such file: No such file or directory

It turns out that I need 32bit libc for cross-compilation which I can’t install since I’m not a superuser. The workaround is to disable this feature by using “--disable-multilib” option as follow:

./configure –-disable-multilib
make

After installation, if you meet the following error in compiling OpenMP file:

gcc: libgomp.spec: No such file or directory,

find libgomp.spec under /path/to/gcc/lib64 and make a symbolic link under /path/to/gcc/lib/gcc/x86_64-unknown-linux-gnu/4.3.2 (See [2])

Reference:
[1] gcc-help mailing list
[2] OpenMP in 30 Minutes

Tuesday, November 04, 2008

Collective Collaborative Tagging System (Abstract)

Currently in the Internet many collaborative tagging sites exist, but there is the need for a service to integrate the data from the multiple sites to form a large and unified set of collaborative data from which users can have more accurate and richer information than from a single site. In our paper, we have proposed a collective collaborative tagging (CCT) service architecture in which both service providers and individual users can merge folksonomy data (in the form of keyword tags) stored in different sources to build a larger, unified repository. We have also examined a range of algorithms that can be applied to different problems in folksonomy analysis and information discovery. These algorithms address several common problems for online systems: searching, getting recommendations, finding communities of similar users, and finding interesting new information by trends. Our contributions are to a) systematically examine the available public algorithms’ application to tag-based folksonomies, and b) to propose a service architecture that can provide these algorithms as online capabilities.

Monday, October 27, 2008

cross-domain problem in making ajax call by using XMLHttpRequest

I’ve just faced with a cross-domain problem in making a mash-up site by using ajax call. In particular, I’m developing an igoogle gadget which should use cross-domain calls to retrieve data from a server outside of google domain.

I’ve found a good article to workaround this problem: http://developer.yahoo.com/javascript/howto-proxy.html

Among the solutions suggested in the article, I’m very pleased with JSONP approaches. Intuition of JSONP approach is that cross-domain problems will not occur for <script> tag. By exploiting this, call a cross-domain function, which designed to return json data enclosed in a function call, by using "src” attribute of <script> tag. Details can be found at http://bob.pythonmac.org/archives/2005/12/05/remote-json-jsonp/

Wednesday, July 30, 2008

Summary of “Evaluation Collaborative Filtering Recommender Systems”

The following is a short summary of:
J. Herlocker, J. Konstan, L. Terveen, and J. Riedl. Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems (TOIS), 22(1):5--53, 2004.

Commonly observed user tasks from recommender systems are as follow:
-. Annotation in context : Uses a recommender in an existing context. E.g., Overlay prediction information on top of existing links
-. Find good items : Video recommender, Item recommender, etc
-. Find All Good Items : purpose is to lower false negative rate
-. Recommend sequence : can be interesting
-. Find credible recommender : making the system more credible

Ratings can be explicit (users rate directly) or implicit (inferred from users behavior or preferences)

Commonly used evaluation metrics are:
1. Predictive accuracy
-. MAE (Mean Absolute Error) = (\sum (p_i – r_i))/N
-. MSE (Mean Square Error) = (\sum (p_i – r_i)^2)/N
-. RMSE (Root Mean Square Error) = \sqrt (MSE)
2. Classification accuracy
-. Precision : the ratio of relevant items selected to number of items selected
-. Recall : the ratio of relevant items selected to total number of relevant items available (Note: Precision and recall are inversely related.)
-. MAP(Mean Average Precision, aka F1) = 2PR/(P+R)

Wednesday, July 16, 2008

Netflix Prize for the best collaborative filtering algorithm

Collaborative filtering, also known as social tagging, is much more popular these days in the Internet but it’s not so clear yet how much and what kind of information we can extract from the system. Regarding this problem, I think Netflix prize would be a great challenge to make those questions clear out.

@. What to predict?

Simple. From the training set, provided by Netflix, which contains over 100M movie 1-to-5 scale ratings, we need to predict unknown movie ratings for the given qualifying set. More specifically, each data in the training set is quadruple of <user, movie, date of grade, grade> and the qualifying set is given <user, movie, date of grade, unknown grade>. We need to fill out the unknown grades by prediction and submit them for evaluation.

Besides two data sets, the training and qualifying set, Netflix provides the probe set which is a problem set with answers. With this, we can roughly estimate the accuracy without consulting with the scoring oracle.

@. How to predict?

Hard. However, we can learn from the front-runners, posted at the Leaderboard. The first annual progress winner is BellKor and they wrote about their algorithm.

@. Number-wise story

#. of data in the training set = 100,480,507
#. of users in the training set = 480,189
#. of movies in the training set = 17,770
#. of data to predict in the qualifying set = 2,817,131
#. of data in the probe set = 1,408,395

@. Research problems

As a student, trying to apply machine learning algorithms to various applications, it is interesting to study:

  1. What kind of machine learning algorithms can be applied to analyze the Netflix collaborative filtering data? Furthermore, Can we find more general algorithms can be used for the data in the net?
  2. How such computations can be expedited by using parallel, multi-core platform? Interestingly, is it possible to use other computing powers, such as cloud computing?
  3. Can we improve accuracy by adding other information easily accessible from the Internet? What kind of infrastructure of the Internet can help this? Web2.0 or else?

@. Reading list

J. Herlocker, J. Konstan, L. Terveen, and J. Riedl. Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems (TOIS), 22(1):5--53, 2004.

R. Bell, Y. Koren, and C. Volinsky. Modeling relationships at multiple scales to improve accuracy of large recommender systems. Proceedings of the 13th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 95--104, 2007

R. Bell and Y. Koren. Improved Neighborhood-based Collaborative Filtering. KDD-Cup and Workshop.

R. M. Bell and Y. Koren, Scalable Collaborative Filtering with Jointly Derived Neighborhood Interpolation Weights, Proc. IEEE International Conference on Data Mining (ICDM'07), 2007

Possible useful papers can be found from the Internet. I will keep this list up-to-date, as I read through.

Tuesday, July 15, 2008

PMPP workshop at NCSA

I attended PMPP (Programming Massively Parallel Processors Agenda) workshop at NCSA last week (July 10th, 2008). It was great opportunity to see what we can do with parallel processors and modern powerful GPUs. Recently Multi-core or many-core has been drawing great attentions in various research communities and people are still struggling to find a way to maximize its capabilities in the coming 80-core era. In other side, Peoples have been also trying to utilize GPUs as a parallel computing unit. Modern GPUs are equipped  with tens of, or hundreds of cores, which can be used for computing intensive jobs. The workshop was about all of these: Multicores and GPUs.

Here are some highlights:

  • Parallel GPU: Mostly two venders, nVidia and ATI, are actively developing this market by continually supplying powerful hardware and SDKs. nVidia provides CUDA and ATI does Stream SDK for programming kits.
  • nVidia's parallel GPU with CUDA: Lightweight threads running on nVidia GPU (8 to 240 cores). Designed with parallel execution in mind, while general CPU is not. GPU can be considered as massively parallel manycore machine.
  • Autotuning for multicore : Multicore tuning parameter spaces are so huge to investigate one by one. We can overcome this by systematic approaches. Details can be seen from the papers by Samuel Williams and Kaushik Datta. In their paper, many multicore tuning techniques are introduced.
  • Cell broadband engine (Cell BE): Interesting demonstration was shown by Hema Reddy from IBM. Look at this article and clip. I saw the more interesting clip in the workshop but I can't find the exact one in the Internet.
  • Multicore/GPU researchers : look at Pedro Trancoso, John Stone, and

Wednesday, April 23, 2008

MDS graph of Connotea

MDS graph with Connotea data by using GGobi.