Showing posts with label science. Show all posts
Showing posts with label science. Show all posts

Tuesday, October 29, 2013

Big data recievers

The overall feeling and initial thoughts when talking about big data is commonly that big data is coming from social media. The most used example of big data is to track what customer think and feel about a company or a product when the write something in social media. This is indeed a very good example and a very usable implementation target for the technical parts that form the basis of handling big data.

However, big data is to be honest much more then social media generated content which can be processed and analyzed. Big data is in general talked about by the four V's. Oracle has identified the 4 V's as Volume, Velocity, Variety and Value while IBM is giving the 4 V's another meaning.

Other sources of big data can be very well the output of the sensors in your factory, sensors in a city, a large number of other sources from within your company or outside of your company. Big data can be the output of devices from the internet of things as described in this blogpost.

Whatever the source of your "big" data the common factors are in most cases that their is a large volume of it, it is coming in a variety of formats and it coming at you fast and continuous. When working with big data the following 3 steps are common;

- Acquire
Their is the need to acquire the data from a number of sources. Within the acquire process their is also the need to store the data. Acquire is not necessarily capturing the data.  In the example of sensor data the sensor will capture the data and send it to the acquire process.

- Process
When data is acquired and stored it will need to be processed. In some (most) cases this will be organizing the data so it can be used in analysis however it can also be very well processed for other means then analysis.

- Use
Using the processed data is in many cases analyzing or further analyzing the data that comes out of the process step. However it can also be input for other, non, analytic business processes.

In the below you can see the vision from Oracle on those steps in which they acquire, organize and analyze data.



Even though the above is giving a good example a step is missing from the visual representation. As it is shown now the acquire phase is displayed as stored data in HDFS, NoSQL or an transactional database. What one of the big parts of a good big data strategy should hold is a step before it is stored and that is receiving the data. One of the big parts where a lot of time will have to be dedicated is creating a good technological capture strategy. You will have to have "listeners" to which data producers can talk and who will write the data to the storage.

If you take for example the internet of things you will have a lot of devices that will be sending out data. You will have to have receivers to which those devices talk. As we have stated that that big data is characterized by volume, variety and velocity this means that you will have to build a number of receivers and those receivers should be able to handle a lot of messages and data. This means that your transceivers should be able to, for example, balance the load and work in parallel to be able to cope with the load of all devices that send data to them. Also you will have to ensure that the write speed of the data that the transceivers need to store the data is in line with the supply of data that is send to the transceivers.

An example of where such topics where an issue and where handled correctly is the CMS DAQ system developed by the TriDAS Project project team at cern when developing data capture triggers for the Large Hadron Collider project. In the below video Tulika Bose a assistant professor from the University of Boston gives a short introduction to this.


Sunday, September 02, 2012

Developing big-data triggers

Most of you will know CERN from the LHC, Large Hadron Collider,  experiment used to discover the Higgs Boson particle. This is one of the most interesting experiments within physics at this moment and the search for the Higgs Boson particle comes into the news quite often. What a lot of people however do not realize is that this research is somewhat different from traditional research in the field of physics as it comes to the amount of data.

When Isaac Newton “discovered” gravity it only took him a tree to lean against and a apple to fall down. Those are not a lot of input streams of information. When it comes to finding the Higgs Boson particle we are playing in a total different field when it comes to the number of inputs. During an event the data capture system will store every second a dataflow the size of rougly six times the Encyclopædia Britannica.

The main issue is that the systems will not be able to handle and store all the data presented to the sensors. All sensors will have triggers developed to capture the most important data. As we are talking about find a particle that is never discovered before the triggers might discard the Higgs Boson particle data instead of storing it for analysis. Developing the triggers is one of the crucial parts of the experiment and one of the most critical parts. In the below video Tulika Bose a assistant professor from the University of Boston gives a short introduction to this.



Within CERN The TriDAS Project is responsible for developing the data Acquisition and High-Level Trigger systems. Those systems will select the data and store it and finally result in data that can be analyzed. For this a large group of scientists and top people from a large number of IT companies have been working together to build this. IT companies like Oracle and Intel have been providing CERN with people and equipment mainly so they can test their new systems in one of the most demanding and data intensive setups currently operational.

Below you can see a high a high level architecture of the CMS DAQ system. This image comes from the "Technical Design Report, Volume 2" delivered by the TriDAS Project project team.

In a somewhat more detailed view the system looks like the architecture below from ALICE project. This shows you the connection to the databases and other parts of the systems.


While finding the Higgs Boson particle is for the common public possibly not that interesting on the short term having IT companies working together with CERN is even though it might not be that obvious at first. CERN is handling a enormous load of data. IT companies who participate in this project are building new hardware, software and algorithms that are specific to finding the Higgs Boson particle. However, the developed technology will be used within building solutions that will end up in serving customers. 

As big-data is getting more and more attention and as we can see all kinds of big-data based solutions are developed we can see that this is no longer a pure scientific play field. It is getting into the day-to-day lives of people. This will help people in the very near future in their day-to-day lives. So, next time you question what the search for the Higgs Boson particle is bringing you as a individual on the short term, take the big-data part into your consideration and do not think it is only interesting to scientists(which is a incorrect statement already however a topic I will not cover on this blog :-) ) 

Saturday, February 11, 2012

String Theory explains why the world ends

Most of us will have quite some difficulty understanding some fields of physics and understanding how all things work. Most of us will already be happy if we understand the basics of some of the leading physics theories and if we can understand bits and pieces of what Einstein was trying to tell us. People like Stephen Hawking and Michio Kaku are able to tell certain parts of their specific fields in such a way that we are understanding the very basics of it. I think that the art of explaining something very complex in a way that the average person understands it is a great gift.

In this video Michio kaku explains the first steps of string theory which is a part of his Floating University lectures. After watching the video the title of this blogpost will be clear to you.

Tuesday, February 07, 2012

Secret Google lab opening doors

Rumors have been going on already for some time on what is behind the website wesolveforx.com . it was rumored some time ago by the New York Times in a article called "Google's lab of Wildest Dreams" that X was referring to a secret lab of Google where they where gathering the brightest minds to solve all kind of issues in the most futuristic ways imaginable. Accoording to the New York Times Sebastian Thrun, one of the world's top robotics and artificial intelligence experts, is a leader at Google Sebastian Thrun, one of the world's top robotics and artificial intelligence experts, is a leader at Google X.

BY looking at the WOIS information of the domain your could see the website was indeed registered by Google. Now Google is releasing more information about "we solve for X" and it turns out that is a great think tank where some great minds are gathering and where everyone else also can participate via the internet to find solutions on some of worlds great challenges in all kind of fields.



Some people already made comparisons to TED.com however in my opinion it go's way future than TED. Within TED you have a broadcasting way of communication where Google X is likely to evolve as a more continuos corporation and social interaction kind of platform. It would be good to keep an eye on the google+ page, the youtube channel and the website of Google X to see where this is heading however it might bring some intersting insights and possibly products in the future.

Or in the words of Google:
"Solve for X is a place to hear and discuss radical technology ideas for solving global problems. Radical in the sense that the solutions could help billions of people. Radical in the sense that the audaciousness of the proposals makes them sound like science fiction. And radical in the sense that there is some real technology breakthrough on the horizon to give us all hope that these ideas could really be brought to life.

This combination of things - a huge problem to solve, a radical solution for solving it, and the breakthrough technology to make it happen - is the essence of a moonshot.

Solve for X is intended to be a forum to encourage and amplify technology-based moonshot thinking and teamwork. This forum started with a small face-to-face event co-hosted by Astro Teller, Megan Smith, and Eric Schmidt - the Solve for X talks are now being posted here on this site. We encourage you to watch the Solve for X talks, join the G+ conversation, and post your own Solve for X talks.
"

Friday, January 20, 2012

Michio Kaku on CERN news

Some scientific experiments get more media coverage then others. When, some time ago, CERN was reporting that during one of their experiments it turned out that neutrinos could travel faster than light, all hell brook loose. For some people it was really strange that some scientists where upset that things could move faster than light, and then another group of people was disappointed. For all who missed some of their classes when we where discussing Einstein and his theories their is good news.

Michio Kaku is explaining why people where upset when they where thinking for a moment that some particles could potentially move faster than light. Michio Kaku is a theoretical physicist, best-selling author, and popularizer of science. He's the co-founder of string field theory (a branch of string theory).

Saturday, December 20, 2008

Partition Decoupling Method

When working on complex systems and trying to map them you will find that the relations within your data can become very complex very fast. Complex data which is time-dependent and interrelated will form a complex maze of data and dependencies which will become almost a impossible task of mapping.

A group of researchers from Darmouth have developed a mathematical tool which can help understand complex data systems like the votes of legislators over their careers, second-by-second activity of the stock market, or levels of oxygenated blood flow in the brain.

“With respect to the equities market we created a map that illustrated a generalized notion of sector and industry, as well as the interactions between them, reflecting the different levels of capital flow, among and between companies, industries, sectors, and so forth,” says Rockmore, the John G. Kemeny Parents Professor of Mathematics and a professor of computer science. “In fact, it is this idea of flow, be it capital, oxygenated blood, or political orientation, that we are capturing.”

Capturing patterns in this so-called ‘flow’ is important to understand the subtle interdependencies among the different components of a complex system. The researchers use the mathematics of a subject called spectral analysis, which is often used to model heat flow on different kinds of geometric surfaces, to analyze the network of correlations. This is combined with statistical learning tools to produce the Partition Decoupling Method (PDM). The PDM discovers regions where the flow circulates more than would be expected at random, collapsing these regions and then creating new networks of sectors as well as residual networks. The result effectively zooms in to obtain detailed analysis of the interrelations as well as zooms out to view the coarse-scale flow at a distance."
Source Press Release


In a paper named "Topological structures in the equities market network" written by Gregory Leibon, Scott D. Pauls, Danile Rockmore and Robert Savell the Partition Decoupling Method is used to map the underlying structure of the equities market network.

"We present a new method for the decomposition of complex systems given a correlation network structure which yields scale-dependent geometric information — which in turn provides a multiscale decomposition of the underlying data elements. The PDM generalizes traditional multi-scalar clustering methods by exposing multiple partitions of clustered entities. "




More information can be found at:
http://www.sciencedaily.com/releases/2008/12/081216131022.htm
http://www.dartmouth.edu/~news/releases/2008/12/16.html
http://arxiv.org/pdf/0805.3470

Sunday, March 16, 2008

Adaptive Algorithms for Online Optimization

Google released a new google techtalk video and now on the subject of Adaptive Algorithms for Online Optimization presented by Seshadhri Comandur a Research Scientist. A list of his publications can be found at his website.

ABSTRACT

The online learning framework captures a wide variety of learning problems. The setting is as follows - in each round, we have to choose a point from some fixed convex domain. Then, we are presented a convex loss function, according to which we incur a loss. The loss over T rounds is simply the sum of all the losses. The aim of most online learning algorithm is to minimize *regret* : the difference of the algorithm's loss and the loss of the best fixed decision in hindsight. Unfortunately, in situations where the loss function may vary a lot, the regret is not a good measure of performance. We define *adaptive regret*, a notion that is a much better measure of how well our algorithm is adapting to the changing loss functions. We provide a procedure that converts any standard low-regret algorithm to one that provides low adaptive regret. We use an interesting mix of techniques, and use streaming ideas to make our algorithm efficient. This technique can be applied in many scenarios, such as portfolio management, online shortest paths, and the tree update problem, to name a few.



Wednesday, October 03, 2007

Women Scared Away From Math?


Are Women Being Scared Away From Math, Science, And Engineering Fields? Have you ever felt outnumbered? Like there are just not that many people like you around? We’ve all felt outnumbered in one situation or another and walking into a situation in which you sense the possibility of being ostracized or isolated can be quite threatening.

One group that may experience this kind of threat is women who participate in math, science, and engineering (MSE) settings- settings in which the gender ratio is approximately 3 men to every 1 woman. Recently, in the wake of comments made by former Harvard University President, Larry Summers, suggesting that women may not possess the same “innate ability” or “natural ability” in these fields as do men, several leading scientific institutions and university presidents publicly lamented the underrepresentation of women in Math, Science and Engineering fields and put out a call to study the reasons for the numbers gap in these areas.

While previous research offers biological and socialization explanations for differences in the performance and representation of men and women in these fields, Stanford psychologists, Mary Murphy and Claude Steele argue that the organization of Math, Science and Engineering environments themselves plays a significant role in contributing to this gap. Murphy contends that situational cues (i.e. being outnumbered) may contribute to a decrease in women’s performance expectations, as well as their actual performance.

Murphy and colleagues showed a group of advanced MSE undergraduates a gender balanced or unbalanced video depicting a potential MSE summer leadership conference. To assess identity threat, the researchers measured the participant’s physiological arousal during the video, cognitive vigilance, sense of belonging and desire to participate in the conference.

The results are telling. The women who watched the gender unbalanced video- where women were outnumbered by men in a 3 to 1 ratio- experienced faster heart rates, higher skin conductance (sweating), and reported a lower sense of belonging and less desire to participate in the conference.

They also found that women were more vigilant to their physical environment when they watched the video in which women were outnumbered. Throughout the testing room, Murphy planted cues related to Math, Science, and Engineering such as magazines like Science, Scientific American, and Nature on the coffee table and a portrait of Einstein and the periodic table on the walls. Women were able to recall more details about the video and the test room, indicating that they paid more attention to the identity-relevant items in order to assess the likelihood of encountering identity threat. “It would not be surprising if the general cognitive functioning of women in the threatening setting was inhibited because of this allocation of attention toward MSE-related cues,” write the authors. Thus, it is likely that this kind of attention allocation would interfere with performance and might help explain the performance gap between men and women in these fields.

While men, in either condition, showed no significant difference in physiological arousal, cognitive vigilance, or sense of belonging, both men and women expressed more desire to attend the conference when the ratio of men to women was balanced. Murphy says that while it’s interesting that both men and women want to be where the women are, the motivations of men and women for wanting to be there are probably quite different. “Women probably feel more identity-safe in the environment where there are more women- they feel that they really could belong there- while men might simply be attracted by the unusual number of women in these settings. Men just aren’t used to seeing that many women in these settings, because the numbers in real Math, Science, and Engineering settings are so unbalanced.”

These findings, which appear in the October issue of Psychological Science, a journal of the Association for Psychological Science, demonstrates that rather than being endemic to women the experience of identity threat in MSE settings is attributable the situation.

This research underlies the importance of situational cues and Murphy hopes that it will "inspire greater motivation to attend to such cues when creating and modifying environments so that they may foster perceptions of identity safety rather than threat."

Note: This story has been adapted from material provided by Association for Psychological Science.

Friday, May 11, 2007

Nasa unveils Hubble's successor

he US space agency Nasa has unveiled a model of a space telescope that scientists say will be able to see to the farthest reaches of the universe.

The James Webb Space Telescope (JWST) is intended to replace the ageing Hubble telescope.

It will be larger than its predecessor, sit farther from Earth and have a giant mirror to enable it to see more.

Officials said the JWST - named after a former Nasa administrator - was on course for launch in June 2013.

The full-scale model is being displayed outside the Smithsonian National Air and Space Museum in the US capital, Washington DC.


The $4.5bn (£2.27bn) telescope will take up a position some 1.5 million km (930,000 miles) from Earth.

It will measure 24m (80ft) long by 12m (40ft) high and incorporate a hexagonal mirror 6.5m (21.3ft) in diameter, almost three times the size of Hubble's.

Hubble, launched in 1990, has sent back pictures of our solar system, distant stars and planets, and remote fledgling galaxies formed not long after the Big Bang.

But scientists say the JWST will enable them to look deeper into space and even further back at the origins of the universe. "Clearly we need a much bigger telescope to go back much further in time to see the very birth of the universe," said Edward Weiler, director of Nasa's Goddard Space Flight Centre.

Martin Mohan of Northrop Grumman, the contractor building the telescope, said that the team was making excellent progress. "There's engineering to do, but invention is done, more than six years ahead of launch," he said.

When ready, the JWST will be launched by a European Ariane V rocket. It is expected to have a 10-year lifespan.Until then, the 17-year-old Hubble telescope will continue to do its work. Nasa plans to send astronauts on the space shuttle to service it in 2008.

JWST is named after James E Webb, Nasa Administrator during the Apollo lunar exploration era; he served from 1961 to 1968 It will be placed 1.5m km from Earth, at Lagrange Point 2, an area of gravitational balance that keeps it in a Sun-Earth line The telescope will be shaded from sunlight by a shield, enabling it to stay cold, increasing its sensitivity to infrared radiation Three principal instruments will gather images of the Universe in the infrared region of the spectrum
These will yield new information about how stars and galaxies first formed a few hundred million years after the Big Bang