Friday, October 10, 2014

Building Boost Libraries quickly

Building Boost

Setting up your development environment

Linux:

1. Fedora/Centos:
You will need:
a. gcc-c++
b. libstdc++
c. libstdc++-devel
Do a yum install or find the rpms on the DVD / net.

2. If you're on *Ubuntu, do a sudo apt-get install of:
a. g++
b. libstdc++-dev
c. libstdc++
Version should be at least 4.8.x - or you'll miss out on the C++11 fun.

3. If you're on Ubuntu and want to try clang, sudo apt-get install the following:
a. clang
b. llvm
c. libc++-dev
Version should be preferably 3.4.

Windows:
If you're on Visual Studio, try to get Visual Studio 2012 or 2013 (Express Edition is free, I have the DVD image for 2013).

Building Boost
Download the boost source archive from boost.org (current version 1.56) and extract it in a directory. Call it <boostsrc>.


Linux:
$ cd <boostsrc>

$ sudo mkdir /opt/boost
$ chown <builduser>:<builduser> /opt/boost
$ ./bootstrap.sh
$ ./b2 install --build-dir=<buildarea> --prefix=/opt/boost --layout=tagged variant=debug,release link=shared runtime-link=shared threading=multi cxxflags="-std=c++11"

<builduser> is the user id of the logged on user.
<buildarea> is the directory where intermediate build products would be generated.

/opt/boost is where the end products would be installed.

Windows:

$ cd <boostsrc>

For 32-bit:
$ "C:\Program Files\Microsoft Visual Studio 12.0\VC\vcvarsall.bat" x86
or
$ "C:\Program Files (x86)\Microsoft Visual Studio 12.0\VC\vcvarsall.bat" x86 


$ b2 install --libdir=<installdir>\libs --includedir=<installdir>\include --build-dir=<buildarea> --layout=tagged variant=debug,release threading=multi link=shared runtime-link=shared

For 64-bit:
$ "C:\Program Files\Microsoft Visual Studio 12.0\VC\vcvarsall.bat" amd64
$ "C:\Program Files\Microsoft Visual Studio 12.0 (x86)\VC\vcvarsall.bat" x86_amd64
$ b2 install --libdir=<installdir>\libs --includedir=<installdir>\include --build-dir=<buildarea> --layout=tagged variant=debug,release threading=multi link=shared runtime-link=shared --address-model=64

<buildarea> is the directory where intermediate build products would be generated.
<installdir> is where you choose to install the products of the build.

Read more!

Wednesday, September 10, 2014

Uses of private inheritance

Found a really cool use of private inheritance in C++ in this answer on stackoverflow: http://stackoverflow.com/a/676725/422131.

More details to follow.

Read more!

Saturday, August 23, 2014

7 useful features in C++14

C++14 is an enhancement over C++11. Here are a few features that are immediately useful:
  1. std::make_unique - a factory function for unique_ptr (which should be the most popular smart pointer) akin to std::make_shared for std::shared_ptr (which shouldn't be the most popular smart pointer).
    #include <memory>
    
    class Foo {
    public:
      Foo(int n) { ... };
      ...
    };
    
    auto fooPtr = std::make_unique<Foo>(10);
    

  2. std::cbegin and std::cend, applied to STL containers give you const_iterators.

  3. std::shared_timed_mutex, a la boost::shared_mutex - a very important addition if you use a fair bit of synchronization. This provides the multiple-reader single-writer (MRSW) type of abstractions.

  4. Getting elements from a tuple by type (if there is a unique element of that type in the tuple).
    std::tuple<int, double, std::string> threeElems = std::make_tuple(1, 2.0, "Foo");
    auto strFoo = std::get<std::string>(threeElems);

    This works, but would have failed had there been two elements of type std::string in the tuple.

  5. Note  how we had to write the type of threeElems in the last example. If we  had used auto instead, its type would be deduced as std::tuple <int, char const*, double> because std::make_shared deduces the type of the  returned tuple using the types of the passed arguments. If you wanted  to use a string literal whose type would be deduced as std::string, you  must use an s suffix like this:
    auto threeElems = std::make_tuple(1, 2.0, "Foo"s);
    auto strFoo = std::get<std::string>(threeElems);
  6. Generic lambdas - essentially lambdas with a very succinct syntax that can be reused for multiple type. Takes a lot of crud away from writing lambdas.
    std::vector<foo> vec;
    std::for_each(vec.begin(), vec.end(), [](auto& elem) { std::cout << elem << '\n'; });

    The  key is the use of the auto keyword for the parameter types. You could  use multiple parameters, all declared auto, whose types are  independently deduced. How does this really help? First, you don't have to write:
    [](Foo& elem) { std::cout << elem << '\n'; }
    
    Also, you could cache a lambda in a generic context with minimal syntactic noise and reuse it for multiple types. Consider this:
    auto elemPrint = [](const auto& elem) { std::cout << elem << '\n'; };
    std::for_each(vecOfInts.begin(), vecOfInts.end(), elemPrint);
    std::for_each(vecOfStrs.begin(), vecOfStrs.end(), elemPrint);
    
    where vecOfInts, vecOfStrs, etc. all contain elements of different, unrelated type.

  7. Being  able to write function with an auto return type, without any trailing  decltype to compute the type. In C++11, you would write something like:
    auto foo(int x, double y) -> decltype(x + y)
    {
       return x + y;
    }
    You can now simply write:
    auto foo(int x, double y)
    {
       return x + y;
    }

    There are several more changes but these stand out in terms of immediate usefulness.

Read more!

Saturday, December 14, 2013

Can you write an assignment operator?

This would be the first in hopefully a series of blog posts to talk of C++11 features and libraries that matter. But in this article, I wouldn't pick up a whole lot of C++11. Instead I shall lay some groundwork first, talking about exception safety in a very informal way and looking at the nothrow swap idiom for copy assignment. Along the way, we'll use some C++11 features (like auto) and libraries (like std::unique_ptr) with obvious syntax and simple usage. Consider a simple class that wraps a character buffer.

#include <iostream>
#include <cstring>
#include <algorithm>

class MyString
{
public:
  // constructor
  explicit MyString(const char *str) : buffer(NULL)
  {
    if (str && str[0] != '\0') {
      auto ln = strlen(str);    // C++11: auto -  
                                //  compiler determines correct type for ln
      buffer = new char[ln + 1];
      std::copy(str, str + ln, buffer);  // more general than strncpy
    }
  }

  // destructor
  ~MyString()
  {
    delete []buffer;
  }

  size_t len() const
  {
    if (buffer) {
      return strlen(buffer);
    } else {
      return 0;
    }
  }

  std::ostream& print(std::ostream& os)
  {
    return (os << buffer);
  }

private:
  char *buffer;
};

What would it mean to copy an object of the above class? What would it mean to assign one object of this class to another? What would the behaviour be of such code:
MyString en("Hello");
MyString es(en);

And of such?
MyString en("Hello");
MyString es("Hola");
en = es;

Without rolling out your own copy constructor and copy assignment operator, disastrous. In the first case the default copy constructor would create object es as a shallow copy of the object en. That would mean that after construction, es and en would both have their data member buffer pointing to the same address. When the scope in which both of these objects are created is exited, es would be destroyed first, followed by en. The destructor of es would have deallocated all the heap-memory pointed to by buffer in one fell swoop, and soon after, en's destructor would try doing the same - and disaster should strike.

The second case is worse in some respects, except that it shouldn't matter: on line 3, as es is assigned to en, the en.buffer starts pointing to the same location as the es.buffer. But en.buffer already pointed to an address at the head of a block of bytes on the heap that had "Hello" in it. Now that both en.buffer and es.buffer point to another location (with "Hola" in it), all references to the "Hello" bytes are lost. This program doesn't have any hopes of being able to track down, and deallocate when it had to, the buffer with "Hello". We have a leak, but it shouldn't matter. Shortly afterwards, when en and es both fall out of scope, es's destructor gets called followed by en's, and as in the case of copy construction above, disaster strikes.

The remedy is well-known - roll out your own copy-constructor and copy-assignment operator.
class MyString
{
public:
  // constructor
  // destructor

  // copy constructor
  MyString(const MyString& that)
  {
    auto ln = that.len();
    if (ln) {
      buffer = new char[ln + 1];
      std::copy(that.buffer, that.buffer + ln, buffer);
    }
  }

  // copy assignment
  MyString& operator = (const MyString& that)
  {
    auto ln = that.len();
    if (this != &that) {
      // release earlier content
      delete [] buffer;
      // and mimic copy construction
      buffer = new char[ln + 1];
      std::copy(that.buffer, that.buffer + ln, buffer);
    }

    return *this;
  }

  // rest of the class
};


Now copy construction creates a copy of the buffer for each new object created and copy assignment takes care of deallocating the older buffer before reallocating the new buffer and copying content. Congratulations. You've just fixed a couple of bad crashes in the code. Bad news, if you wrote this code in an interview, they'll offer you a good C++ book and not the job. Porque? Que pasa? Because you goofed up the copy assignment. For what would happen if the call to std::copy threw? Ok, in this rather unimaginatively contrived example, it would likely not. But in general, we carry out several steps in the assignment: deletion of the old buffer, allocation of a new buffer and then copying. If the allocation of the new buffer fails, you have no way to get back and salvage your older data. Nor if the copy fails after that. In simple terms, the code we've written is not exception safe.

The key problem is losing the previous buffer before the new buffer is ready. If we first create the new buffer separately, then cache the old buffer, assign the new buffer and finally delete the old buffer, we've made a start.
class MyString
{
...
  MyString& operator = (const MyString& that)
  {
    auto ln = that.len();
    if (this != &that) {
      // allocate and set aside
      char *new_buffer = new char[ln + 1];
      std::copy(that.buffer, that.buffer + ln, new_buffer);
      
      // cache the old
      char *old_buf = buffer;
      // assign the new
      buffer = new_buffer;
      // delete the old
      delete [] old_buf;
    }

    return *this;
  }
...
};
But problems still abound. If copy threw, we'd be left with a leak. Besides, we are still dealing with a single member and this scheme quickly gets out of hand if you deal with two or more members with similar requirements. We can make a small improvement here.
class MyString
{
...
  MyString& operator = (const MyString& that)
  {
    if (this != &that) {
      // Use RAII
      MyString tmpStr(that.buffer);
      
      // swap the two pointers
      std::swap(buffer, tmpStr.buffer);
      // et voila!
    }

    return *this;
  }
...
};
If an exception is thrown before line 8, nothing changes. If one is thrown after line 8, tmpStr.buffer is deallocated by a call to its destructor. The call to swap cannot throw. Once that call is complete, ownership of buffers have been exchanged and the destructor of tmpStr takes care of deallocating the older buffer of the current object (this). If we are dealing with multiple members, extending this logic requires a little extra effort. Define a swap member function, or specialize std::swap for MyString, and implement it with no-throw guarantees. A set of pointer swaps for one should be able to provide that guarantee. Your code would then look like:
namespace std
{
  void swap(MyString& lhs, MyString& rhs)
  {
    if (&lhs != &rhs) {
      char *tmp = lhs.buffer;
      lhs.buffer = rhs.buffer;
      rhs.buffer = tmp;
    }
  }
}

class MyString
{
...
  MyString& operator = (const MyString& that)
  {
    if (this != &that) {
      // Use RAII
      MyString tmpStr(that.buffer);
      
      // swap the two objects
      std::swap(*this, tmpStr); // or swap(tmpStr) if swap were a member
      // et voila!
    }

    return *this;
  }
...
};
This is the standard idiom for writing copy assignments using no-throw swaps and on another day I would have happily concluded this article here. Alas! We still have a problem. If you've been attentive you may have already noticed it. What if the MyString constructor threw at line 20? It mighty well can, if say the call to std::copy threw. Ok, I hear you - it won't in the case of this example. But we are performing two operations in the constructor - allocation and assignment of values to the cells of the allocated buffer. If the latter operation throws, the destructor of MyString won't get called and we'd be leaking the memory allocated for buffer. The fool-proof way to deal with the lack of atomicity of this kind of resource allocation plus initialization issues is to harness RAII in some form to protect the smallest units of allocation. We'll use a C++11 smart pointer to do the trick for us. Here is the full listing.
#include <iostream>
#include <cstring>
#include <algorithm>
#include <memory>

class MyString
{
public:
  // constructor
  explicit MyString(const char *str)
  {
    if (str && str[0] != '\0') {
      auto ln = strlen(str);
      buffer.reset(new char[ln + 1]);
      std::copy(str, str + ln, buffer.get());
    }
  }

  // copy constructor
  MyString(const MyString& that)
  {
    auto ln = that.len();
    if (ln) {
      buffer.reset(new char[ln + 1]);
      std::copy(that.buffer.get(), that.buffer.get() + ln, buffer.get());
    }
  }

  // destructor
  ~MyString()
  {}

  size_t len() const
  {
    if (buffer) {
      return strlen(buffer.get());
    } else {
      return 0;
    }
  }

  // copy assignment
  MyString& operator = (const MyString& that)
  {
    if (this != &that) {
      // copy the right side
      MyString tmp(that);

      // relinquish our data's ownership
      // to tmp, and acquire tmp's data
      swap(tmp);
    }

    return *this;
    // let tmp go out of scope and release
    // our older data in its destructor
  }

  // nothrow swap
  void swap(MyString& rhs)
  {
    buffer.swap(rhs.buffer);
  }

  std::ostream& print(std::ostream& os)
  {
    return (os << buffer.get());
  }

private:
  // C++11 smart pointer to make resource
  // management of buffer exception-safe
  std::unique_ptr<char[]> buffer;
};


int main()
{
  MyString m1("Hello"), m2("Hola");
  MyString m3(m1);
  m1 = m2;

  m1.print(std::cout) << std::endl;
  m2.print(std::cout) << std::endl;
  m3.print(std::cout) << std::endl;
}

Three points to note:
  • The member buffer is now a std::unique_ptr smart pointer (actually its std::unique_ptr specialization for arrays).
  • If std::copy throws on line 15 or 25 in the constructor, the destructor of buffer is called correctly and there is no leak.
  • std::unique_ptr provides a no-throw swap function which can be used to perform the copy assignment.
Prior to C++11 standard library smart pointers were limited to auto_ptr and they would be of limited use here. Using unique_ptr from C++11 makes the code a whole lot succinct. To be sure, the only real difference between the last listing and the one before that is in how we wrapped individual units of allocation (buffer) in RAII wrappers (std::unique_ptr). This last listing can be seamlessly extended to more such members and would still work.

Read more!

Sunday, October 21, 2012

Enforcing source code formatting guidelines

Last few days there has been an uptick in interest in programming processes at the workplace. Folks believe code-reviews need to be taken more seriously and we should revisit all the written guidelines which have existed for eons. Such guidelines include a relatively elaborate and fairly well-written coding standard among other things.

I spent the second half of last week creating a checklist for code review and in the process came up against all sorts of emotions about coding standards. These ranged from utter neutrality and mild intolerance to rabid disgust. The common refrain on guidelines about checking indentation, padding, etc was to give them the least priority. Unfortunately code reviews are done by most people in one pass and they have to filter each line through all the standard concerns of the reviewer - from correct functionality to correct padding and indentation. One of my colleagues suggested if we could use the IDE for enforcing formatting. It set me off on a wild-goose hunt for cool scripts / plugins for C++ on vim. I managed to find one (google.vim) and customized it a fair bit but figured that it was good for indentation and that's about it.

After grappling with the vim scripting syntax and trying out some obscenely brute-force approaches to formatting code via vim (involving vim scripting, unreadable regexps and all sorts of command-chaining) I figured I wasn't even half-way through. A little googling finally brought up a couple of code formatters: uncrustify and Artistic Style. I finally picked up the latter because of the former's documentation poverty. All impressed with astyle and for good reason:
  1. Offers a fairly granular set of options to pad expressions, indent statements, place brackets, add or remove blank lines, etc.
  2. Is very simple to use. Here's how you format a source file from the command line:
    $ astyle proj/src/lib/mysource.cpp
  3. Has a concise but really useful documentation.

Ok, #2 doesn't just work like that. You have to either create a file called .astylerc in your home directory (on Unix) or use set of command-line switches. I'd recommend the first approach. Here is a sample .astylerc file.

One would typically run astyle on the entire code base once and commit it to the repository. Later on it should be run on each source file each time it is committed to the repository. The only point of discomfort with tools of this type is the fact that they edit your code to fix indentation, padding and other formatting issues. I would much rather have a tool like Google's cpplint.py which points out issues but leaves it to the user to fix them. However the style enforced by cpplint.py is hard-coded to follow Google's own recommendations which differs in some matters from what we follow in our organization. So somebody has to read and edit cpplint.py to suit our purposes.

Astyle works on Windows I but haven't tried it yet. But it works perfectly on Linux and I am impressed with how little I had to try to create a .astylerc file that enforces our coding standard. After this very no one should complain about the reviewer fussing over whitespace.

Read more!

Thursday, November 04, 2010

Return your objects by value (at least sometimes)

Starting where we left a year and a half back (and I swear I had no time for the blog in this intervening period), we'll look at temporaries again, but in a slightly different light. We saw in that article that it's a great idea to eliminate temporaries of non-POD types (or more correctly, types with non-trivial copy semantics) as far as possible. One big advantage of eliminating temporaries is the elimination of redundant copying from a temporary to a named instance which is how a lot of temporaries inevitably end up. Common ways in which temporaries get created are when objects are returned by value from a function call, and also when a function (say foo) is invoked in-place in the argument list of the invocation of another function (say bar) - and thus the return value of foo() is passed to bar(). Some of these cases are illustrated in the following code:

Fud foo(int i) // Fud is a copiable class
{
return Fud(i); // Temporary created
}
void bar(Fud);
...
bar(foo(1)); // Temporary passed to bar
Fud f = foo(2);
...

Now suppose we were to rewrite the above code to eliminate temporaries. There are a few different ways, but let's try a fairly simple approach:

Fud foo(int i) // Fud is a copiable class
{
return Fud(i); // Temporary created
}
void bar(const Fud&);
...
bar(foo(1)); // const-reference to temporary, no copying
Fud f = foo(2);
...

In the above code, we have eliminated the pass by value of a Fud object to bar(...) but foo(...) still returns a temporary by value. Here is an enhancement:


void foo(int i, Fud *&f) // Fud is a copiable class
{
f = new Fud(i);
}
void bar(const Fud&);
...
Fud *pfud = NULL;
foo(2, pfud);
bar(*pfud); // const-reference to temporary, no copying
...

This time there don't seem to be any temporaries and consequently no redundant copying. But the code is no longer simple. At the least, you cannot do something like bar(foo(...)) any more - nesting calls is a natural algebraic operation used to compose functions but the elimination of references takes that ability away from us. Strictly speaking, with Boost shared_ptr, we could eliminate this limitation (how? left as an exercise to the reader, an author's exclusive prerogative). But it still means that we have to incur the cost of dynamic memory allocation. Again, dynamic memory allocation is not necessarily evil - sometimes, it is even more welcome than allocating a large stack based object. But the overall readability of code has suffered as well.
What do we do? Well, it depends.
A good idea is to start by passing temporaries back by value. Now don't cross your eyes - what you just read is exactly what I said and there is a little something that the C++ standard allows (without mandating) and most standard compilers implement as an optimization, which makes this possible. It is called Copy Elision and in a special form, Return Value Optimization or RVO for short.

Copy Elision


Copy elision simply refers to elimination of redundant copying, if at all possible. For example, consider the following code:

Fud obj = Fud(1);

What's happening there? A Fud instance called obj is initialized from a temporary Fud object created through Fud(1). Normally, we would expect that Fud's constructor will be invoked to create a Fud object with an initialization parameter value of 1. Next, the instance called obj will be created and initialized through a call to its copy-constructor which will copy the state of the temporary object to obj. However, it is not hard to see that the only purpose of the temporary object is to help initialize the obj instance. This code could have been written in a much simpler way as:

Fud obj(1);

There would have been no calls necessary to the copy constructor of obj, nor would there be any temporary instance to destruct. The behaviour of the code would have been exactly the same (unless of course the copy constructor or destructor had side-effects - always a bad idea). Copy elision is the inbuilt optimization in the compiler which causes code like this:

Fud obj = Fud(1);

to generate the equivalent of code like this:

Fud obj(1);

and thus prevent needless copying. An interesting special case is Return Value Optimization (RVO). It is best illustrated with the following example:

#include <vector>
#include <string>
#include <iostream>

using std::vector;
using std::string;
using std::cout;
using std::endl;

struct TestClass
{
TestClass(int i) : i_(i)
{}

TestClass(const TestClass& tc) : i_(tc.i_)
{
cout << "Copied." << endl;
}

private:
int i_;
};

vector<TestClass> make_vec()
{
vector<TestClass> vec;
vec.reserve(20);

cout << "Starting populating vector." << endl;
for (int i = 0; i < 10; i++) {
vec.push_back(TestClass(i));
}
cout << "Vector populated." << endl;

return vec;
}

int main()
{
vector<TestClass> vec = make_vec();
}

Look at line 40 above. The function make_vec() returns a vector which is copied to the vector vec - what looks like copy initialization. Since the contained type of the vector is TestClass, you'd expect the TestClass instances to be copied, first time when they are enqueued in the vector inside make_vec (line 31) and again when the returned temporary vector from make_vec is used for copy-initialization of vec at line 40. That would be conformant behaviour, but that's often not the behaviour you get. Running on gcc 4.3.2 on OpenSuSE gave me this output:

Starting populating vector.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Vector populated.

This clearly shows that there is no copying during copy-initialization of vec - in other words there is no copy-initialization. This is an example of Copy Elision - a special case known as RVO. Internally, the compiler might generate code that arranges for a reference to vec on line 40 to be passed to the function make_vec, and obviate the local vector inside make_vec so that all push_back operations happen on this reference instead. A slight change to the above code can completely disable Copy Elision / RVO, as illustrated below:

return vec;
}

int main()
{
vector<TestClass> vec;
vec = make_vec();
}

I got this:

Starting populating vector.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Vector populated.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.
Copied.

Quite clearly, the copy-initialization takes place and has not been optimized away.

The big advantage of RVO is that the syntax of function calls can be made to match the semantics of the mapping that the function represents. The return value rather than an out-parameter is a return value. This also allows for nested function calls, a natural algebraic expression of functional composition.

Therefore, depending on how the return value of function is going to be used - it might be a good idea to return an object by value.

Read more!

Monday, February 02, 2009

Jargon Time: L-values, R-values and Temporaries

As promised here are a few fairly basic examples of C++ jargon, demystified, here in this article. We start looking at basics of C++ expressions: l-values, r-values and temporaries. The concepts are simple, but important, and would be used in later columns of this (Jargon) series.

L-values and R-values


When we write code, we express action and intent through well-formed expressions that conform to a broad syntax. Some of this action involves moving data around, some of it involves carrying out a more complex operation - and most involve both. For example:

Show line numbers
 double number = 0.0;
number = 2.0;
double square_root = ::sqrt(number);

The above code involves both moving data around, and carrying out some action. In the second line, the literal double 2.0 is assigned to the double variable number. In this context number is an l-value expression - because it allows modification of the value it holds, when used on the left hand side of an assignment expression.

The expression 2.0 on the other hand can only be used to assign values to expressions such as number - in other words it can only be used on the right hand side of assignment expressions, never on the left hand side. It can also be used in function return statements. Such expressions are called r-values.

At this point we have three small observations to make:
1. An r-value can never be used on the left hand side of an assignment operation.
2. An l-value in very many cases can be used on the right hand side of assignment operations also. In this case, it simply degenerates to the value it contains. For example:

Show line numbers
 double number = 0.0;
number = 2.0;
double anotherNumber = number;

In the above code snippet, on the third line, the expression number is used as an r-value and degenerates to the value contained in the l-value expression number.
Some l-values cannot be used as r-values. For example, in the above code snippet, the expression double anotherNumber is an l-value expression, but we cannot write code like:

double aThirdNumber = (double anotherNumber = 2.0);

Language rules do not allow this. So you have an example of an l-value expression that cannot degenerate to an r-value expression.
3. l-values can also be used in function return statements. However, whether it is treated as an l-value or simply degenerates to an r-value depends on the return type of the function. If a function returns a non-const reference or pointer to an object, the function call can be considered as an l-value expression. That's essentially because the return value of the function is an l-value. For example:

Show line numbers


template<int size>
struct CheckedIntArray {
int& operator[](int index) {
if (index >=0 && index < size) {
return array_[index];
}
throw IndexOutOfBoundsException; // some exception
}
private:
int array_[size];
};

In the above CheckedIntArray class, operator[](int) can actually be used as an l-value expression because it returns a reference to an element in the underlying array. This enables use to write code like this:
Show line numbers
 CheckedIntArray<16> my_array;
my_array[0] = 15;


In the above, the expression my_array[0] = 15; is equivalent to my_array.operator[](0) = 15;.

Many complex expressions are r-values. As well as a few simple ones:

Show line numbers
 int i = 0;
++i; // r-value expression

Arrays are r-values although individual elements in an array are not. Of course this does not apply to a pointer being used with an array syntax. Thus:

Show line numbers
 int arr[32] = {0};
int arr2[32] = {1};
arr[0] = 5; // arr[0] is an l-value
// the following is illegal
// arr = arr2; // arr is an r-value



Finally, let it be said that all expressions in C++ are either l-value or r-value expressions.

Temporaries


Related to the concept of r-values is the concept of temporaries. In fact temporaries are r-values (without all r-values being temporaries). Consider the following example:

Show line numbers
 int m = 4;
int n = 5 + 8/m;

Here the expression 5 + 8/m is a temporary. This is a relatively simple temporary - possibly one that would only exist in the registers of the CPU. However, it is possible, and quite common, to have temporaries on the stack. The important thing to understand is that temporaries are unnamed values, which are created in the context of an expression and whose life time is limited to the period of evaluation of that expression. Consider the following expression:

string str = string("Hola amigos!");

The right hand side expression creates a temporary string object, and it is then copied to a local variable called str. Once the control of the executing program reaches past the semi-colon terminating this line of code, the temporary object is gone. Only str, containing a copy of it, exists.

There is one exception to this rule and it deals with references. In the last expression, if instead of a string variable on the left, we had a string reference, things would be a little different:

const string& str = string("Hola amigos!");

First of all, if you see we've had to add a const to the reference. We could not have had a non-const reference to a temporary. This is always the case, as you can see below:

Show line numbers
 const int& r = 5;
const double& s = 2.0;

Since all temporaries are r-values, it is clear a non-const reference cannot refer to them. But, the exception that I referred to is in the life time of the temporary when a (const) reference refers to it. In this case, the temporary persists till the reference is in scope, and not just till the end of the statement that created the reference.

References are often created for function return values, although most optimizing compilers would eliminate the creation of these temporaries if the return value of the function was not assigned to any specific object. In general, reducing the number of temporaries that is created by a program is a good strategy for optimization, and to some extent, the compiler already does it.

Since all temporaries are r-values, it is clear a non-const reference cannot refer to them. But, the exception that I referred to is in the life time of the temporary when a (const) reference refers to it. In this case, the temporary persists till the reference is in scope, and not just till the end of the statement that created the reference.

References are often created for function return values, although most optimizing compilers would eliminate the creation of these temporaries if the return value of the function was not assigned to any specific object. In general, reducing the number of temporaries that is created by a program is a good strategy for optimization, and to some extent, the compiler already does it.

As a final example of how temporaries are generated, and where we can run into trouble with them if we are not careful, I present a piece of code I have seen written in several places (including products I have worked on).
Show line numbers
 using std::string;
using std::stringstream;
using std::cout;
using std::endl;

...

int x = 0;
double f = 1.6;
stringstream sout;
sout << "Some data values streamed: " << x << "|" << f;
const char *str = sout.str().c_str();
cout << str << endl; // this will likely print garbage

Can you spot the trouble with the above code. The trouble is that the member function std::string str() const of the std::stringstream class returns a temporary string. But in the expression sout.str().c_str(), we get a reference to the const char* pointer member of the returned temporary string and we copy it to the variable called str. As soon as this statement is executed, the temporary that was created as a result of the call to sout.str() is destroyed. But we still have a dangling pointer referring to its internal char * string, which is invalid for all good money. Needless to say, the last line above can even crash the program itself.

In the next edition of the Jargon column, we'll look at Namespace lookups and the Interface Principle. Keep watching, for more jargons demystified.

Read more!